Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

389 results about "Image transformation" patented technology

Thyroid tumor diagnosis method and system based on ultrasonic and cytological image conjoint analysis

The invention discloses a thyroid tumor diagnosis method and system based on ultrasonic and cytological image conjoint analysis. The method comprises the following steps: converting a thyroid B ultrasonic image of a patient to be diagnosed into an image feature vector FUS; performing structured feature extraction on the thyroid cytological image of the patient to be diagnosed to obtain a structured feature vector FCYTO; the image feature vector FUS and the structured feature vector FCYTO are converted into a fusion Token sequence; and inputting the Token into a multi-modal feature fusion and prediction network based on a Transform architecture, and carrying out feature fusion and classification prediction so as to obtain the probability that the thyroid tumor of the patient to be diagnosed is malignant. The thyroid tumor diagnosis based on multi-modal fusion is carried out on the basis of ultrasonic and cytological images, so that the diagnosis accuracy is improved.
Owner:金凤实验室

Multi-scale linear array camera splicing method and system based on point cloud

The invention discloses a multi-scale linear array camera splicing method and system based on point cloud. The method comprises the following steps: completing acquisition and preprocessing of point cloud data and image data of a target area; determining an overlapping region range between adjacent images; extracting spatial structure characteristics in the point cloud data, and performing multi-scale hierarchical decomposition on the point cloud through a multi-scale segmentation method; meanwhile, multi-scale image feature extraction is carried out on the images of the linear array camera; solving gradients in X and Y directions by adopting an optical flow method aiming at any pixel in the overlapping region, and calculating a motion vector between the pixels; fusing the optical flow information obtained under each scale, and constructing a globally consistent optical flow vector field; according to the fused optical flow vector, calculating to obtain a geometric transformation matrix of the whole overlapping region; and after image transformation and alignment are completed through the transformation matrix, fusion processing is carried out on overlapped areas. And the unification of the visual effect and the spatial integrity of the spliced image is ensured.
Owner:WUHAN HANNING TECH

Image contour recognition method and device for cell mass

The invention discloses an image contour recognition method and device for a cell cluster, and relates to the technical field of biological image processing, and the method comprises the steps: converting an input cell cluster image into a binary image; carrying out expansion operation and distance transformation on the image after morphological operation, and carrying out subtraction operation on the two obtained images to obtain a third image containing an unknown region; detecting connected regions in the accurate foreground region, and marking each independent connected region, background region and unknown region to obtain a marked image; performing region segmentation on the original image to obtain segmented marked regions, generating a corresponding binary mask image, and extracting the contour of each marked region by using the binary mask image; identifying and filtering the contour of each marked area to obtain an effective contour; the size of each effective profile is converted into an equivalent sphere diameter. According to the invention, accurate segmentation of adjacent or adhesion areas in the image can be efficiently realized, and accurate counting and size measurement of cell clusters can be obtained.
Owner:BEIJING ESSENTIA BIOSCIENCES LTD

Image defogging enhancement and intelligent perception collaborative optimization method based on stream matching

The invention provides an image defogging enhancement and intelligent perception collaborative optimization method based on stream matching, and the method comprises the steps: obtaining a foggy input image, and constructing an ordinary differential equation for defining image transformation; the input image is input into a fog perception vector field, the fog perception vector field comprises an atmospheric scattering purifier and a defogging perception color lookup table, and the atmospheric scattering purifier is used for extracting multi-scale features from the foggy input image to generate clearer output; the defogging perception color lookup table adaptively adjusts the image color through a nonlinear color transformation mechanism; and solving the ordinary differential equation through an RK4 solver, iteratively updating the foggy input image at each time step, and ensuring that the input image can be stably converted into a clear image from an initial foggy state. The technical problems that in the prior art, the existing method is forced to reduce the model capacity due to the calculation efficiency constraint in the ultra-high-definition image defogging, so that the fog concentration estimation is incomplete, and the perception quality and the calculation efficiency cannot be considered at the same time are solved.
Owner:SUN YAT SEN UNIV

Teaching student network for end-to-end semi-supervised object detection

A system and method for end-to-end semi-supervised object detection is provided. The system retrieves labeled and unlabeled images from an image dataset and generates an input batch by application of image transformation(s) on the images. The system further generates a first result for each image of the input batch by application of a teacher neural network on the input batch. For an object in an unlabeled image of the batch, the first result includes candidate bounding boxes and scores for the boxes. The system determines a threshold score based on the scores and selects a foreground bounding box from the candidates. The system generates a second result by application of a student neural network on the unlabeled image and computes a training loss over the input batch based on the foreground bounding box and the second result. The system trains the student neural network based on the training loss.
Owner:SONY GROUP CORP

Biological image transformation using machine-learning models

Described are systems and methods for training a machine-learning model to generate image of biological samples, and systems and methods for generating enhanced images of biological samples. The method for training a machine-learning model to generate images of biological samples may include obtaining a plurality of training images comprising a training image of a first type, and a training image of a second type. The method may also include generating, based on the training image of the first type, a plurality of wavelet coefficients using the machine-learning model; generating, based on the plurality of wavelet coefficients, a synthetic image of the second type; comparing the synthetic image of the second type with the training image of the second type; and updating the machine-learning model based on the comparison.
Owner:INSITRO INC

Systematic testing of AI image recognition

Disclosed are systems and methods including software processes for developing test cases for testing robustness of AI-based image-recognition models-under-test (MUTs) with respect to types of image variation transformations. The system may generate various types of robustness metrics for the MUT and output user-readable reports about the MUT's performance. The system trains machine-learning architectures to generate test cases including augmented images according to the types of image transformations, applies the IR MUTs, and then evaluates the image feature vector embeddings and predicted classification produced by the IR MUTs to determine the accuracy of the MUT with respect to each type of transformation.
Owner:FRAUNHOFER USA INC

Biological image transformation using machine-learning models

Described are systems and methods for training a machine-learning model to generate image of biological samples, and systems and methods for generating enhanced images of biological samples. The method for training a machine-learning model to generate images of biological samples may include obtaining a plurality of training images comprising a training image of a first type, and a training image of a second type. The method may also include generating, based on the training image of the first type, a plurality of wavelet coefficients using the machine-learning model; generating, based on the plurality of wavelet coefficients, a synthetic image of the second type; comparing the synthetic image of the second type with the training image of the second type; and updating the machine-learning model based on the comparison.
Owner:INSITRO INC

Panoramic image determination method and device, equipment, medium and product

The invention discloses a panoramic image determination method and device, equipment, a medium and a product. The method comprises the following steps: acquiring a current scene image, shot by at least one camera device, of an environment to which a vehicle belongs, and kinematics data and attitude data of the vehicle from a current shooting moment to a previous shooting moment; determining a scene image transformation matrix based on the kinematics data and the attitude data; acquiring a historical scene image shot by the camera device at a previous shooting moment, and performing stereo matching on the current scene image and the historical scene image based on the scene image transformation matrix and shooting parameters of the camera device to obtain a scene stereo image corresponding to the camera device; and determining and displaying a panoramic image based on the scene stereo image of each camera device. The authenticity, accuracy and reliability of panoramic image determination are improved, and the driving safety is improved.
Owner:CHINA FAW CO LTD

Progressive cross-view-angle image geographic positioning method based on spatial feature aggregation and position perception

The invention relates to the technical field of image geographic positioning, in particular to a progressive cross-view-angle image geographic positioning method based on spatial feature aggregation and position awareness, which comprises the following steps: acquiring paired satellite images and ground panoramic images as sample images, and constructing a training sample set; a cross-view image geographic positioning model composed of a double-branch backbone network and a fine-grained prediction network is constructed, two branches of the double-branch backbone network are a satellite processing branch and a ground panorama processing branch, and the fine-grained prediction network is composed of a bird's-eye view image transformation module and a position sensing prediction module; inputting training samples in the training sample set into the cross-view image geographic positioning model, and performing iterative training on the cross-view image geographic positioning model until convergence; and inputting a to-be-positioned paired satellite image and ground panoramic image into the trained cross-view-angle image geographic positioning model to obtain a positioning result output by the trained cross-view-angle image geographic positioning model.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Automatic image cutting method and system based on face key points

The invention provides an automatic image cutting method and system based on face key points, which are applied to the technical field of image processing, and the method comprises the steps: inputting a portrait image into a portrait segmentation model for foreground extraction, and obtaining a figure region mask; performing top positioning estimation based on the figure area mask to obtain a top position of the portrait image; inputting the portrait image into a face key point detection model to obtain a plurality of key points; determining a face midpoint of the portrait image based on the position relationship of the plurality of key points; performing affine transformation estimation based on the source point set and a preset target point set to obtain an affine transformation matrix; adjusting the vertical offset of the affine transformation matrix to obtain a longitudinal correction matrix; performing edge correction on the longitudinal correction matrix to obtain a cutting transformation matrix; and performing image transformation on the portrait image based on the cutting transformation matrix to obtain a standard composition image. According to the invention, the cut image with uniform size, standard composition and good detail retention can be generated.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI +1

Location-aware text search and visualization capabilities for physical environments

A computing system may include an image access engine configured to access a panoramic point cloud image of a physical environment. The computing system may also include an environment location-aware text engine configured to transform the panoramic point cloud image into an alternate representation that reduces distortion in the panoramic point cloud image and perform an optical character recognition (OCR) process on the alternate representation to determine text in the panoramic point cloud image. The environment location-aware text engine may further be configured to construct text labels to track the text determined in the panoramic point cloud image and support text searches for the physical environment through the text labels.
Owner:SIEMENS INDUSTRY SOFTWARE INC

View-conditioned diffusion for real-world vehicle gaussian splatting

Systems and methods for view-conditioned diffusion for real-world vehicle gaussian splatting. A single perspective image can be transformed using image transformation techniques to generate a training dataset that addresses a domain gap between synthetic data and real-world data in a traffic scene. A pre-trained diffusion model can be finetuned with the training dataset to obtain a fine-tuned diffusion model. Perspective-aware images having different perspective views of an entity from the single perspective image can be generated using the fine-tuned diffusion model. A large generative model (LGM) can be trained using the perspective-aware images to generate a gaussian splatting model for the entity. View-conditioned simulations from the single perspective image can be generated by using the gaussian splatting model for downstream tasks.
Owner:NEC LABORATORIES AMERICA INC

Similarity optimal on-board image registration method based on inertial navigation data fast convergence

The application discloses a similarity optimal on-board image registration method based on inertial navigation data fast convergence, and comprises the following steps: constructing a projection transformation model between an image transformation matrix, a to-be-registered image and a reference image; obtaining the projection transformation matrix according to the field of view optical axis position and three-angle offset of the satellite at the imaging moment and coarse registration; performing global coarse matching by using the projection transformation model and the projection transformation matrix to obtain a coarse matching image; decoupling the image transformation matrix according to a mathematical model to obtain a to-be-solved transformation parameter, substituting the to-be-solved transformation parameter into the image transformation matrix to obtain an optimal transformation matrix, and applying the optimal transformation matrix to the to-be-registered image to obtain a fine registration image; and calculating the similarity of the fine registration image and the reference image according to a local normalized image similarity measurement algorithm to obtain an optimal matching result. The application has the advantages of meeting the on-board calculation capacity requirement, meeting the image processing precision requirement and meeting the on-board storage resource requirement.
Owner:BEIJING RES INST OF SPATIAL MECHANICAL & ELECTRICAL TECH

System and method for automatically testing confused APP based on visual identification

The invention relates to the field of visual identification, and discloses a visual identification-based confused APP automatic test system. The system comprises an image input module, an image preprocessing module, a manifold learning module, an information geometry matching module, an image conversion module, a dynamic element processing module, an automatic operation execution module and a report generation module. The invention further provides an automatic confused APP testing method based on visual identification, and the method comprises the steps of accurately positioning a control and executing automatic testing operation through technologies such as image preprocessing, manifold learning, information geometric calculation and Kalman filtering, and finally recording a result and generating a report. By combining manifold learning, information geometry, Kalman filtering and other technologies, the control positioning precision and the dynamic element adaptive capacity are improved, the image preprocessing efficiency is optimized, automatic test operation execution is achieved, the problems of inaccurate control recognition and the like in the prior art are solved, and the test stability and reliability are improved.
Owner:BEIJING COMANDA ELECTRONIC TECH CO LTD

Knowledge distillation method and system based on view alignment

The invention relates to the technical field of machine learning, in particular to a knowledge distillation method and system based on view alignment. The invention aims to solve the main limitations in the traditional distillation technology, including the problems of excessive confidence of a teacher model, confirmation deviation and the like. The invention provides a novel knowledge distillation framework based on logarithmic probability, which is called KDVA; specifically, z-score standardization is applied to the logarithmic probability of a model for smooth output, so that transmission of more teacher implicit knowledge is promoted. In addition, a same-view-angle and cross-view-angle alignment mechanism is introduced into the KDVA, and comparison information of weak and strong image transformation is utilized to enlighten a student model to obtain more knowledge. In addition, the student model is supervised by using the real label of the sample, so that the student model can obtain more information about each sample target category. According to the method, when different network architectures are applied to different data sets, high effectiveness and stability are shown.
Owner:BOZHOU UNIV

Image generation system, image generation method, and program

An image generation system includes a first image acquirer, a second image acquirer, and an image processor. The image processor performs image transformation processing and superposition processing. The image transformation processing includes generating, based on a first image, a transformed image by subjecting a predetermined extracted part of a first object to image processing. The image transformation processing includes: processing of dividing the extracted part into a plurality of segments; and processing of changing at least one parameter selected from the group consisting of a grayscale, a location, a size, and an orientation of at least one segment belonging to the plurality of segments.
Owner:PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD

Mask conditioned image transformation based on a text prompt

In accordance with the described techniques, an image transformation system receives an input image and a text prompt, and leverages a generator network to edit the input image based on the text prompt. The generator network includes a plurality of layers configured to perform respective edits. A plurality of masks are generated based on the text prompt that define local edit regions, respectively, of the input image for respective layers of the generator network. Further, the generator network generates an edited image by editing the input image based on the plurality of masks, the respective edits of the respective layers, and the text prompt.
Owner:ADOBE INC

A semi-supervised dual-network medical image segmentation method

The present invention discloses a semi-supervised dual-network medical image segmentation method, comprising the following steps: processing a data set, randomly dividing the data set into a training set and a test set in a 4:1 ratio, and performing image enhancement; building a DNSS network model, using a geometry-aware network and an image transformation consistency-based network, and adding an uncertainty perception module at the end of the auxiliary network to effectively utilize unlabeled data; training the DNSS network on the training set, performing the segmentation task and generating a segmentation model; testing the model on the test set, and selecting the model with the best performance as the final model based on the test results and saving it. The present invention can generate efficient and accurate labeled data, reduce the reliance of medical image segmentation on manually recorded data, reduce the cost of manually labeling medical images, and assist doctors in diagnosis.
Owner:JIANGSU UNIV OF SCI & TECH

Vision-based robot vascular suture force feedback estimation method, medium and system

The invention belongs to the field of medical robots, and particularly relates to a robot blood vessel suturing force feedback estimation method, medium and system based on vision, and the method comprises the following steps: collecting and preprocessing a high-frequency dynamic image of a blood vessel region and a suturing tool; performing time-frequency analysis on the image; after the image is converted to a frequency domain, a motion mode of the cardiovascular system is extracted; constructing a mechanical model based on the extracted motion mode and a prior biomechanical model of the cardiovascular system; and combining the extracted motion modal information with the mechanical response of the mechanical model through a space-time information fusion algorithm, and estimating the contact force when the suture tool is in contact with the blood vessel. Compared with the prior art, the problem that a force feedback estimation method based on vision is difficult to apply in a dynamic environment in the prior art is solved. According to the scheme, more accurate force control and operation are achieved, and the safety and the success rate of the robot vascular suturing operation can be improved.
Owner:RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE

Passive domain adaptive three-dimensional medical image segmentation method based on continuity constraint and difficulty guidance

The invention provides a continuously constrained and difficulty guided passive domain adaptive three-dimensional medical image segmentation method. The method comprises the following steps: in a source domain pre-training stage, carrying out full-supervised training on a segmentation model by utilizing source domain annotation data; in the pseudo source domain image generation stage, a thought of combining coarse generation and fine generation is adopted, style migration is performed by using a frozen source domain pre-training segmentation model and target domain unlabeled data in coarse generation, and a target domain image is converted into a pseudo source domain image with a source domain style; and in the fine generation step, Fourier transform is utilized to remove artifacts and noise in the coarsely generated image. In the target domain adaptation stage, a pre-training segmentation model, a pseudo source domain image and a target domain image are utilized, and continuity constraint between slices and a difficult sample mining mechanism are fused to carry out an adaptation process from a source domain to a target domain. According to the method, under the condition that source domain data does not need to be accessed, the spatial context constraint and the difficult sample mining mechanism of the three-dimensional medical image are effectively fused.
Owner:FUZHOU UNIV

Infrared image registration method and unsupervised learning image registration model training method

The invention discloses an infrared image registration method and an unsupervised learning image registration model training method, and belongs to the technical field of image processing. Aiming at the characteristics of weak texture and few features of a low-overlapping-rate infrared image, a weak texture feature extractor is firstly designed, and the feature extraction capability of a model on a weak texture image is improved through multi-scale feature fusion; then, by adopting multi-scale receptive field correlation calculation, the discrimination capability of the model on similar features is enhanced; furthermore, by adopting a homography transformation estimation method which is optimized gradually from coarse to fine, the estimation precision of a transformation matrix is improved; and finally, realizing registration and splicing of the infrared images through image transformation and overlapping region fusion processing. According to the method, the registration precision of the low-overlapping-rate infrared image can be effectively improved, and application scenes such as large-scene image splicing can be supported.
Owner:国网湖北省电力有限公司直流公司

Infrared vision positioning method for concrete pumping liquid level in tube of steel tubular arch bridge

An infrared vision positioning method for concrete pumping liquid level in tube of steel tubular arch bridge is provided. In this method, the perspective geometry algorithm is used to construct the pose transformation matrix to obtain the relative pose between the camera and the arch bridge steel pipe, the image is transformed into orthography by perspective transformation, and the real coordinates of the liquid level in the infrared image can be directly solved by the scale, so as to realize the accurate positioning of the liquid level; the proposed method is simple in structure and can realize real-time operation, which greatly improves the calculation efficiency of coordinates in infrared images, compared with the conventional knocking method, the advantage is that it does not require staff in an overhead operating, and the feedback of the liquid level position can be faster and more efficient.
Owner:CHONGQING JIAOTONG UNIV

IMU (Inertial Measurement Unit)-based real-time video image stabilization method and system for intra-frame motion compensation

The invention provides a real-time video image stabilization method and system for performing intra-frame motion compensation based on an IMU (Inertial Measurement Unit). The method comprises the following steps: S1, preprocessing data; s2, track smoothing processing is carried out; s3, performing intra-frame motion compensation: S3.1, initializing line exposure delay time; s3.2, calculating the time difference between the current vth line and the middle line exposure in the current frame image; s3.3, acquiring the smoothed virtual camera pose of the current frame and the real camera pose of each layer; s3.4, acquiring a motion compensation matrix for compensating the line of image to the virtual pose of the frame of image; s3.5, acquiring an image frame perspective transformation matrix according to the camera internal reference matrix K and the motion compensation matrix; s3.6, performing intra-frame motion compensation on the image layers to obtain an initial motion compensation matrix queue; s3.7, carrying out black edge removal processing on the motion compensation matrix queue; obtaining a motion compensation matrix queue after black edge removal; and S4, image transformation output. The system comprises a data preprocessing module, a track smoothing module, an intra-frame motion compensation module and an image transformation output module. The operation efficiency is improved, and the anti-shake real-time performance is improved.
Owner:INGENIC SEMICON CO LTD

Method for deploying a robotic system to scan inventory within a store based on local wireless connectivity

One variation of a method for deploying a mobile robotic system to scan inventory structures within a store includes: dispatching the mobile robotic system to navigate along inventory structures within the store during a setup cycle; at the mobile robotic system, while navigating along the inventory structures during the setup cycle, capturing a set of wireless connectivity metrics representing connectivity to a first wireless network; assembling the set of wireless connectivity metrics into a wireless connectivity map of the store; estimating a processing duration from start of the scan cycle to transformation of images of the inventory structures, captured by the mobile robotic system, into a stock condition of the store; and dispatching the mobile robotic system to autonomously capture images of the inventory structures within the store during a scan cycle preceding a scheduled restocking period in the store based on the processing duration.
Owner:SIMBE ROBOTICS INC

Identifying and localizing editorial changes to images utilizing deep learning

The present disclosure relates to systems, methods, and non-transitory computer readable media that utilize deep learning to identify regions of an image that have been editorially modified. For example, the image comparison system includes a deep image comparator model that compares a pair of images and localizes regions that have been editorially manipulated relative to an original or trusted image. More specifically, the deep image comparator model generates and surfaces visual indications of the location of such editorial changes on the modified image. The deep image comparator model is robust and ignores discrepancies due to benign image transformations that commonly occur during electronic image distribution. The image comparison system optionally includes an image retrieval model utilizes a visual search embedding that is robust to minor manipulations or benign modifications of images. The image retrieval model utilizes a visual search embedding for an image to robustly identify near duplicate images.
Owner:ADOBE INC +1

Method and mobility device for generating aligned image data through aligning parameters generated by an image transformation artificial intelligence model

A method for generating aligned image data through an aligning parameter generated by an image transformation artificial intelligence (AI) model includes, through an encoder of the image transformation AI model, generating at least one or more aligning parameters from a first camera property and a second camera property related to a first camera and a second camera respectively. The method also includes, through an image transformer of the image transformation AI model, transforming, based on the at least one aligning parameter and a brightness parameter, first image data photographed by the first camera to be aligned with second image data photographed by the second camera. The method also includes training the encoder and a discriminator of the image transformation AI model by adversarial training. The image transformation AI model discriminates between the transformed first image data and the second image data.
Owner:HYUNDAI MOTOR CO LTD +1

System and method for image correction in a camera system using adaptive image deformation

A video meeting system for adjusting a perspective using adaptive image morphing includes an image morphing unit including at least one processor. The at least one processor is programmed to: receive an overview video stream from a camera in a video meeting system; determining at least one region of interest represented within the at least one test frame based on an analysis of the at least one test frame from the overview video stream; determining one or more indicators of actual camera perspective relative to the at least one region of interest; determining a target camera perspective relative to the at least one region of interest, the target camera perspective being different from the actual camera perspective; determining at least one image transformation based on a difference between the actual camera perspective and the target camera perspective; applying at least one image transformation to one or more sub-frame regions of the plurality of image frames of the overview stream to generate at least one image-warped primary video stream; and causing the at least one image-warped primary video stream to be displayed on a display.
Owner:HUDDLY INC

Adaptive document integration using generative artificial intelligence

Systems, methods, and computer-readable media are provided for using generative AI enriched with metadata about historical document characteristics to transform documents of various formats, including images, to the fields and values they represent. A prompt template may be selected in association with a type of document. The prompt template indicates field definition(s) of field(s) to be detected in the document and location(s) in which the field(s) have been detected in prior documents. A large language model is prompted with a prompt generated using the prompt template to generate a result that assigns value(s) to the field(s). Output from the language model is used for identifying the field to value mapping for the document, such that data detected from the document may be stored in appropriate database structures of a database. Metadata stored in association with the prompt template is updated based on location(s) in the document in which the field(s) were detected, and the value(s) of the field(s) are stored in a database. Outbound documents may be similarly translated to detect values of corresponding fields requested by third parties, even if those values are not stored in the database. In this scenario, values for fields may be detected in outbound documents using the prompt templates enriched with metadata as processed by the large language model before such information is prepared to be sent to a third party.
Owner:ORACLE INT CORP