Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

583 results about "Image code" patented technology

Mathematical problem solving method and device based on multi-modal large model and electronic equipment

The invention provides a mathematical problem solving method and device based on a multi-modal large model and electronic equipment, and relates to the technical field of artificial intelligence. The method comprises the following steps: determining a mathematical element image of a mathematical problem; inputting the mathematical element image into an image coding model to obtain an image vector output by the image coding model; the image coding model is obtained by training based on a sample mathematical element image and a positive sample text description and a negative sample text description corresponding to the sample mathematical element image; inputting the image vector into a self-adaptive module to obtain an image conversion coding vector output by the self-adaptive module; the adaptive module is obtained based on training of a sample image vector and a sample text vector; determining question stem characters of the mathematical question, and inputting the question stem characters and the image conversion coding vector into the large language model to obtain a prediction answering process output by the large language model; the large language model is obtained based on sample question stem characters, sample image conversion coding vectors and sample answering process training, and the mathematical problem solving capability of the multi-modal large model can be improved.
Owner:TSINGHUA UNIVERSITY

Remote sensing target detection method and device based on diffusion model

The invention relates to the technical field of target detection, and provides a remote sensing target detection method and device based on a diffusion model. The method comprises the following steps: step 1, acquiring a remote sensing image data set, and preprocessing the remote sensing image data set; 2, training an image coding module based on the preprocessed remote sensing image data set; the image coding module is used for extracting multi-scale features of a remote sensing image to obtain a multi-scale feature map; 3, constructing an initial noise frame through a variable variance scheduling method, and training a detection decoding module based on the initial noise frame and the multi-scale feature map extracted by the image coding module; the detection decoding module performs noise prediction on the initial noise frame and recovers a target frame to complete target detection; and 4, forming a remote sensing target detection model by using the trained image coding module and the detection decoding module, wherein the remote sensing target detection model is used for performing target detection on the remote sensing image.
Owner:HENAN UNIVERSITY

Optical flow estimation method and system fusing Mama and visual basis model knowledge

The invention belongs to the technical field of computer vision and deep learning, and particularly relates to an optical flow estimation method and system fusing Mama and visual basic model knowledge. The method comprises the following steps: performing down-sampling feature extraction on two adjacent frames of input images by using a convolutional neural network to obtain local texture features; performing down-sampling on the first frame image to obtain context features; meanwhile, extracting global semantic features of two adjacent frames of images by using a pre-trained visual model, and performing adaptive fusion enhancement through an adaptive semantic texture feature fusion module to obtain an image coding feature pair after semantic enhancement; constructing a related volume through pixel-by-pixel dot product operation; and finally, based on the obtained related volume and context features, iteratively optimizing the output optical flow through a loop iteration updating module. The method solves the problems that in a low-texture, repeated-texture or sheltered area, feature expression is unstable, self-adaptive modeling capacity is lacked, different scenes are difficult to generalize, and model performance and efficiency cannot be balanced.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Information processing method and system based on cellular pixel structure

The invention discloses an information processing method and system based on a honeycomb pixel structure, and the method comprises the steps: an encoding step: encoding a positioning graph and input information to form an encoding matrix based on regular hexagonal pixels, the encoding matrix comprising a positioning graph region and an effective information region; a masking step, at least performing XOR operation on the effective information area of the coding matrix and different mask matrixes to form different cellular codes, and comparing the different cellular codes to obtain an optimal cellular code; the mask matrix is also a matrix based on hexagonal pixels; and a decoding step: decoding the optimal cellular code to obtain effective information. According to the method, the cellular codes can be quickly, accurately and efficiently coded and decoded, and a new choice is provided for application of an image coding technology.
Owner:ZHEJIANG DAOMING OPTOELECTRONICS TECH

Model training method, electronic equipment and computer readable storage medium

The invention discloses a model training method, electronic equipment and a computer readable storage medium, and relates to the technical field of data processing and large models. The method comprises the steps of obtaining a training data set, wherein the training data set comprises image annotation data of various types of concerned objects and description text annotation data associated with the image annotation data; training the initial multi-modal retrieval model by adopting the training data set to generate a learnable prompt; carrying out image coding training on the initial multi-modal retrieval model by adopting the training data set and the learnable prompt to generate an intermediate multi-modal retrieval model; the training data set is adopted to conduct text and image alignment training on the middle multi-modal retrieval model, a target multi-modal retrieval model is generated, and the target multi-modal retrieval model is used for conducting multi-modal retrieval on the target retrieval request to obtain a target retrieval result. According to the method and the device, the technical problems that a multi-modal retrieval model in related technologies cannot adapt to variability and large-scale data requirements, and the model performance is relatively poor are solved.
Owner:ALIBABA CLOUD FEITIAN (HANGZHOU) CLOUD COMPUTING TECH CO LTD

Image coding method and system based on single-step diffusion model

The invention provides an image coding method and system based on a single-step diffusion model, and the method comprises the steps: coding a potential representation to a decoding end through employing an extremely low bit rate based on a stable diffusion model (SD-Turbo) and a depth potential representation compression model, and obtaining a reconstructed image through single-step denoising; a group of auxiliary encoders and decoders are introduced, rich and entropy-perceived pixel-level original image semantic information is extracted at an encoding end, and decoded image structure information is shared for a single-step denoising process at a decoding end; and the whole model is subjected to end-to-end optimization by adopting code rate and pixel level constraint so as to achieve the optimal subjective coding quality. According to the invention, the technical problems of high decoding complexity and poor reconstruction consistency are solved.
Owner:STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST

Triangulation-based adaptive subsampling of dense motion vector fields

The present disclosure relates to an apparatus and a method for providing a plurality of motion vectors related to an image coded in a bitstream. The method includes obtaining a set of sample positions within the image, obtaining respective motion vectors associated with the set of sample positions, deriving an additional motion vector based on information coded in the bitstream, determining an additional sample position located within a triangle, which is formed by three sample positions of the set of sample positions, based on the respective motion vectors associated with the three sample positions, the triangle not including the rest of the sample positions of the set of sample positions, adding the additional sample position to the set of sample positions, and associating the derived additional motion vector with the additional sample position.
Owner:HUAWEI TECH CO LTD +1

Automobile headlamp communication system and method based on image frame insertion

The invention relates to the technical field of optical communication, in particular to an automobile headlamp communication system and method based on image frame interpolation, comprising a transmitting end and a receiving end, a vehicle-mounted central control system of the transmitting end is used for receiving input to-be-transmitted information, and an ambient light detection module evaluates an external illumination condition in response to the input of the to-be-transmitted information; an image gray value and frame insertion frequency are set based on the illumination intensity, the image coding module codes information to be transmitted based on the set image gray value to generate a projection image frame, and the projection module inserts the projection image frame into a normal illumination frame sequence of a vehicle lamp based on the frame insertion frequency and projects the projection image frame to an external environment; the image acquisition module of the receiving end is used for acquiring a projection image frame, the image decoding module preprocesses the projection image frame and then decodes the projection image frame, and the information to be transmitted is restored and sent to the vehicle-mounted central control system. According to the invention, vehicle lamp communication at a certain distance in a traffic environment can be realized on the premise of ensuring normal vehicle lamp illumination.
Owner:SUZHOU UNIV

Image coding prediction mode selection method

The invention discloses an image coding prediction mode selection method, which solves the problems that in the prior art, image redundant information is not efficiently utilized in conventional image coding, and complex video coding is relatively high in application delay for low hardware cost, and the like. The absolute reference sample value and the recovery sample value of the reference image are used as reference sample values, so that the image redundant information is effectively utilized, and the compression efficiency is improved; a fixed number of coding prediction modes are set in a selection set, different coding prediction mode selection sets are determined for different scenes, and coding code streams are reduced; when the optimal coding prediction mode is selected, resource consumption and delay are reduced by means of source image information, so that high coding efficiency and good subjective effect can be realized under the condition of low cost and low delay.
Owner:NANJING MAGEWELL ELECTRONICS CO LTD

Water hammer pressure prediction and active suppression method and system

The invention relates to the field of fluid pipeline system safety and intelligent control, and provides a water hammer pressure prediction and active suppression method and system. The method comprises the following steps: acquiring a multi-dimensional digital signal flow, and generating a window data matrix sequence through sliding window segmentation and dimensionless; image coding is carried out to form a spatiotemporal feature sequence, the spatiotemporal feature sequence is processed by a double-branch spatiotemporal prediction network, and a prediction feature spectrum sequence is output and decoded to form a prediction waveform matrix; future risk indexes and morphological characteristics are extracted and spliced with the multi-dimensional digital signal flow to form a prediction enhancement state vector, and the prediction enhancement state vector is input into a strategy network to generate valve action control parameters; and a target valve opening curve is decoded, a composite control signal is obtained through feedforward-feedback composite control, actual response data flow and historical data packaging are collected, and a prediction network and a strategy network are updated online through mixed priority screening. According to the method, prospective and self-adaptive water hammer prediction and active suppression are realized, the pressure evolution trend of a pipeline system is informed, and decision support is provided for safe operation of a fluid pipe network.
Owner:SHANGHAI MONDIAL TEST & ASSEMBLY SYST INC

Video generation method based on mask with body

The invention discloses a video generation method based on a body mask, belongs to the technical field of artificial intelligence, and can solve the problems that an existing body world model is inconsistent in an action space and a pixel space, is sensitive to the change of a visual angle of a camera, and is not uniform in architecture among different body structures. The method comprises the following steps: S1, determining a mask sequence with a body according to a target video; s2, encoding the body mask sequence and the initial frame image of the target video respectively to correspondingly obtain body mask features and image encoding features; s3, inputting the body mask features into a control network module to obtain injection features, and inputting the image coding features into a backbone network of a video generation model to obtain backbone layer features; and S4, fusing the injection feature and the trunk layer feature to obtain a fused feature, and generating a prediction video according to the fused feature. The method is used for generating the predictive video with the body.
Owner:ZHONGKE FIFTH CENTURY (HANGZHOU) INTELLIGENT TECHNOLOGY CO LTD +1

Image recognition method, apparatus and device, storage medium, and product

Embodiments of the present application provide an image recognition method, apparatus and device, a storage medium, and a product. According to the technical solution provided in the embodiments of the present application, an image content recognition model uses an image-text interaction unit to perform analysis processing on a target text feature and a target image feature of an image to be recognized, to obtain a content recognition result of said image; moreover, the target text feature is generated on the basis of a set target review rule by means of a text coding unit in the image content recognition model, and the target image feature is generated on the basis of said image by means of an image coding unit in the image content recognition model. Image-text interaction can be effectively performed on the basis of the target text feature of the target review rule and the target image feature of said image, thereby accurately detecting metaphorical information in the image, and effectively improving the image recognition effect and efficiency.
Owner:BIGO TECH PTE LTD +1

Industrial design-oriented drawing semantic analysis and structured conversion method and system

The invention discloses an industrial design-oriented drawing semantic analysis and structured conversion method and system. The method comprises the following steps: acquiring an input image, carrying out different-scale coding on image features through an image coding backbone network, selecting a specific layer to extract a multi-scale feature map, and fusing to generate multi-scale features; inputting the multi-scale features into a geometric primitive detection network and an other element detection network, respectively identifying geometric primitives and non-geometric primitives, and carrying out positioning and classification; analyzing mutual relations between geometric elements and other elements contained in the graph through an element relation graph network, and generating formalized language description based on a matching rule; a Prompt template is constructed, geometric information is supplemented by using a multi-modal large model, the rationality of the supplemented information is verified through a geometric attribute relationship verifier, and complete formalized language description is obtained. According to the method, the industrial drawing image containing the complex constraint relation can be efficiently and accurately converted into machine-readable structured data.
Owner:XI AN JIAOTONG UNIV

Video coding method, video decoding method, video coding device, video decoding device and storage medium

The invention provides a video coding method and device, a video decoding method and device and a storage medium, and the video coding method comprises the steps: obtaining a video frame sequence; for any video frame in the video frame sequence, calculating an inter-frame optical flow change index of the video frame; the inter-frame optical flow change index represents the motion change degree between the video frame and the previous video frame; dividing the video frame sequence into a key frame set and a redundant frame set; performing image coding on each key frame in the key frame set to obtain an image code stream; performing semantic text recognition on each redundant frame in the redundant frame set through various recognition models, and generating a semantic text code stream according to a plurality of semantic text recognition results corresponding to the plurality of redundant frames one by one; and sending the image code stream and the semantic text code stream to a cloud server. According to the embodiment of the invention, different downstream tasks can be flexibly adapted by extracting information of different dimensions.
Owner:PEKING UNIV

Method for training image-text matching model, computing device, and storage medium

A computer-implemented method is provided. The method includes: obtaining a sample text and a sample image corresponding to the sample text; labeling a true semantic tag for the sample text according to a first preset rule; obtaining a text feature representation of the sample text and a predicted semantic tag output by a text coding sub-model; obtaining an image feature representation of the sample image output by an image coding sub-model; calculating a first loss based on the true semantic tag and the predicted semantic tag; calculating a contrast loss based on the text feature representation of the sample text and the image feature representation of the sample image; adjusting parameters of the text coding sub-model based on the first loss and the contrast loss; and adjusting parameters of the image coding sub-model based on the contrast loss.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Camera condition guide viewpoint synthesis method based on video diffusion model

The invention relates to a video diffusion model-based camera condition-guided viewpoint synthesis method, which belongs to the technical field of computer vision, and comprises the following steps of: taking a stable video diffusion model SVD as a basic generator, and jointly inputting Gaussian noise and encoded image potential representation to obtain a video diffusion model SVD-based viewpoint synthesis model; meanwhile, a pose encoder based on a time attention mechanism is used for encoding camera parameters represented by Plcarbon coordinates, and then the camera parameters are used as pose conditions to be embedded into a time attention layer in a video diffusion model de-noising U-Net; pixel-level features of an input single image are extracted through an image coding embedding network and serve as image feature conditions to be embedded into a space attention layer in a de-noised U-Net of a video diffusion model, and a pre-trained video diffusion model is finely adjusted; the space consistency of the synthetic image and the input image and the track consistency of the synthetic image and the camera parameters are improved, and the generation diversity is increased while the generation quality is improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Multimodal assembly action recognition method for comparing semantic query

The invention discloses a multi-modal assembly action recognition method for comparative semantic query, and relates to the technical field of man-machine cooperation assemblation.The method comprises the steps that a visual sensor is arranged on an assembly workbench to obtain an operator action video, a sampling frame sequence is obtained through random frame sampling, a skeleton sequence is obtained through human body posture estimation, and a skeleton sequence is obtained through human body posture estimation; and inputting an assembly action recognition model to complete recognition. The model comprises an image coding module, a skeleton coding module, a feature fusion module, a text coding module and a semantic comparison module which are used for extracting image and skeleton features, fusing features, coding preset category text description, comparing action features with category text features and outputting a result with the highest similarity, and a comparison loss function is adopted during training. According to the method, multi-modal information is fused, the problems of single-modal limitation and multi-modal semantic segmentation are solved, category text semantics are fully utilized, the fine-grained action recognition precision is improved, the over-fitting risk is reduced, and the generalization and task migration ability of the model in a dynamic industrial scene is enhanced.
Owner:ZHEJIANG UNIV

Urban environment performance prediction method and system based on multi-modal fusion and attention enhancement

The invention provides an urban environment performance prediction method and system based on multi-modal fusion and attention enhancement, and the method comprises the steps: collecting image data and numerical data, and generating a plurality of urban environment performance distribution truth value maps; preprocessing the multi-modal data to obtain an image map and a numerical value feature vector, and pairing the image map and the numerical value feature vector with the true value map to form a multi-modal data set; a multi-target model of a conditional generative adversarial network based on attention enhancement is constructed, a generator of the multi-target model comprises a two-way encoder, spatial features can be extracted based on an image map, and physical features can be extracted based on a numerical feature vector; an image coding path adopts a U-Net down-sampling structure containing a convolution block attention module, and a numerical value coding path adopts a multi-layer perceptron; and the multi-head output layer outputs a plurality of predicted urban environment performance distribution diagrams. After the model is trained, target area data are input, and three types of prediction distribution diagrams are output. According to the invention, the urban environment performance prediction effect is improved.
Owner:HUNAN ARCHITECTURAL DESIGN INST +1

Road extraction method

The invention provides a road extraction method in order to solve the problem that extracted roads are not connected due to the fact that an existing large-scale road extraction method based on remote sensing images is insufficient in discontinuous road structure modeling capacity and lack of display topology supervision. According to the invention, a model construction method based on a wave particle dipictorial view angle is adopted, particle features and volatility features in an image are extracted respectively, the volatility features output by a fluctuation encoder and the particle features output by a particle encoder are fused through a feature interaction module, the road structure detail and semantic expression ability is enhanced, and the image quality is improved. The method solves the problem that the road features are submerged or neglected in the extraction stage, and achieves the screening and full storage of image coding information. A geometric perception multi-constraint loss function designed for road topological structure features is adopted to perform supervised learning training on the node extraction network, so that the problem of weak supervision pertinence of the node extraction network in a training process in the road extraction network is effectively solved, and the accuracy of road extraction is improved.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Visual coding method, visual coding model training method and device

The invention provides a visual coding method, a visual coding model training method and a visual coding model training device in the field of computer vision, which are used for realizing coding of images with different resolutions by using the same visual coding model and are applied to image coding scenes with more sizes. Comprising the steps that firstly, an input image is acquired, and the input image can be a high-resolution image or a low-resolution image; and then inputting the input image into a visual coding model for dividing the input image into a plurality of image blocks according to position embedding, extracting features from each image block, and coding and outputting the visual coding data based on the features of each image block and the corresponding position, the position embedding is obtained by adjusting the initial position embedding according to the difference between the input image and the preset resolution, and the position embedding can specifically comprise a corresponding matrix when the input image is divided.
Owner:HUAWEI TECH CO LTD

Personalized face image generation method and system

The invention relates to the technical field of personalized images, and provides a personalized face image generation method and system, and the method comprises the steps: inputting a reference image and a preset text prompt set into a double-layer text coding module, and obtaining a first text prompt containing a subject term vector s *; inputting the text prompt into a pre-trained stable diffusion model, and training a double-layer text coding module by using LDM loss; inputting the reference image into a double-layer face image coding module to obtain an image prompt, and inputting the reference image into a trained double-layer text coding module to obtain a second text prompt; inputting the image prompt and the second text prompt into a stable diffusion model through a decoupling cross attention mechanism, and training a double-layer face image coding module by using LDM loss; and prompting and inputting the reference image and the personalized prompt text into a stable diffusion model containing a trained double-layer text coding module and a double-layer face image coding module to generate a personalized face image.
Owner:SUN YAT SEN UNIV

Crack detection method and system

The invention discloses a crack detection method and system, and belongs to the technical field of crack detection. The method comprises the following steps: acquiring to-be-detected crack data; inputting the to-be-detected crack data into a pre-constructed crack detection model, and outputting a crack detection result; the crack detection model comprises an image coding module and a multi-source learnable prompt mechanism module which are arranged in parallel, and a mask decoding module connected behind the image coding module and the multi-source learnable prompt mechanism module, a rough channel edge sensing module is further connected behind the image coding module, and the output of the mask decoding module and the output of the rough channel edge sensing module are connected through a grading multiplication module. The rough channel edge sensing module is introduced into the crack detection model, multi-level features are effectively integrated, the sensing ability of crack boundaries is enhanced, and background interference is inhibited; a multi-source learnable prompt mechanism module is designed, semantic perception guidance is dynamically generated, deep fusion and adaptive enhancement of multi-level edge features can be realized, and morphological features and boundary details of fine cracks are effectively captured.
Owner:HOHAI UNIV +2

Image coding method and device, storage medium and program product

The invention discloses an image coding method and device, a storage medium and a program product, and relates to the technical field of image compression, and the image coding method comprises the steps: processing an original image through a perceptual quantization network, and determining an image token corresponding to each sub-block in the original image; generating text information corresponding to the original image through a cross-modal mapping technology and the image tokens; and performing context compression on the text information through a preset large language model to generate a code stream corresponding to the original image. According to the invention, through the context association capability of the perceptual quantization network, the cross-modal mapping and the large language model, the code rate of the code stream is reduced, the compression rate of image compression is improved, and the image is convenient to store and transmit.
Owner:PEKING UNIV SHENZHEN GRADUATE SCHOOL

Scene text super-resolution method based on multi-domain prior guidance

The invention discloses a scene text super-resolution method based on multi-domain prior guidance, and the method comprises the following steps: obtaining a multi-scene text image, carrying out the data preprocessing, and constructing a data set; training a pre-constructed neural network model based on the data set to obtain a corresponding scene text super-resolution model; the scene text super-resolution model comprises a text prior enhancement branch and a super-resolution branch fused with text prior; the text prior enhancement branch comprises a frequency domain perception text image coding module, a multi-domain prior registration module and a multi-domain prior interaction registration module; and inputting a scene text image needing super-resolution processing into the trained scene text super-resolution model to obtain a corresponding scene text super-resolution image. According to the method, spatial domain and frequency domain feature information is combined, multi-domain prior guidance generated through interaction is utilized, and a super-resolution text image which is high in resolution and easy to read is guided to be generated.
Owner:NO 15 INST OF CHINA ELECTRONICS TECH GRP

Ultrasonic image boundary perception segmentation method, system and device

The invention discloses an ultrasonic image boundary perception segmentation method, system and device, and relates to the field of image processing, and the method comprises the steps: obtaining an ultrasonic image, and sequentially processing the ultrasonic image through a block embedding module and an image coding module to obtain spatial domain features; performing frequency domain feature extraction processing on the spatial domain feature by using a preset multi-scale frequency extraction strategy to obtain a fine-grained high-frequency feature, a coarse-grained high-frequency feature and a low-frequency feature, and obtaining a frequency fusion feature in combination with a preset frequency alignment strategy; a preset frequency guide boundary refining strategy is combined to determine refining features; and according to the refined features and a preset boundary guiding decoding strategy, determining a segmentation mask representing a result of segmenting the target region under the ultrasonic image from the background region. A low-frequency structure and multi-scale high-frequency boundary details are explicitly separated through frequency domain decomposition, an ultrasonic image boundary sensing segmentation scheme based on frequency guidance is provided, the boundary sensing ability is high, and the cross-domain generalization ability is high.
Owner:THE UNIV OF NOTTINGHAM NINGBO CHINA

Image processing method and device, electronic equipment and computer readable storage medium

The invention relates to the technical field of image processing and the field of insurance business and smart medical treatment, and provides an image processing method and device, electronic equipment and a computer readable storage medium, and the method comprises the steps: obtaining a to-be-processed image; encoding the to-be-processed image to obtain a plurality of image block features; performing identification processing on the to-be-processed image based on the region detection model, determining a key region image, and constructing a key patch set according to the key region image; performing calculation processing on the information entropy of each image block feature to obtain an image information entropy; determining a compression weight value based on the activation function and the image information entropy; determining weighted compression image feature information based on the image block features and the corresponding compression weight values; and constructing a final image feature vector by using the key patch set and the weighted compressed image feature information. According to the technical scheme, the image coding calculation cost can be reduced, and then the reasoning rate of the visual language model can be well improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Dynamic image based contactless instructions for an electronic monitoring system

An electronic monitoring system for secure image code generation and a method of activating devices in the electronic monitoring system via the secure image codes are provided. The electronic monitoring system includes a base station in communication with a monitoring device. The monitoring device includes a camera to obtain images of a monitored area. A controller for the electronic monitoring system receives a request for a desired interaction with the electronic monitoring system from a first device and generates multiple image codes corresponding to the desired interaction. The controller transmits the image codes to a mobile device and then receives each of the image codes from the camera when the image codes are displayed on the mobile device while the mobile device is positioned in front of the camera. The controller activates the desired interaction with the electronic monitoring system responsive to receiving the image codes from the camera.
Owner:ARLO TECHNOLOGIES INC

Image coding method and apparatus, image decoding method and apparatus, readable medium, and electronic device

An image coding method and apparatus, an image decoding method and apparatus, a readable medium, and an electronic device are disclosed. The image coding method includes: obtaining an original image, and performing block processing to obtain a plurality of image blocks; calculating a gradient value of a pixel in each image patch, and screening for important region blocks according to the gradient values of the pixels; and inputting the important region patches and position information of the important region patches in the original image into a visual conversion model for coding so as to generate a bit stream.
Owner:CHINA TELECOM CORP LTD

Electron microscope microscopic image dislocation instance segmentation method based on visual large model

The invention relates to a dislocation instance segmentation method of an electron microscope microscopic image, and belongs to the technical field of material defect detection and image processing. Aiming at the problems of low efficiency and poor precision of traditional manual analysis, the invention provides a multi-scale dislocation instance segmentation method based on a visual large model, the method comprises an image coding module, a multi-scale feature generation module and a mask decoding module, a multi-scale feature pyramid is generated by fusing global and local features, model training is optimized in combination with a parameter efficient fine tuning technology, and a multi-scale dislocation instance segmentation result is obtained. And bounding box positioning, category classification and pixel-level mask segmentation of dislocation instances are realized. The method remarkably improves the detection efficiency and precision, supports automatic material quality evaluation, reduces personal errors, and can be widely applied to dislocation defect analysis in the fields of semiconductors, aerospace and the like.
Owner:CHONGQING UNIV