Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

86 results about "Visual segmentation" patented technology

Flexible grabbing method of robot

The invention discloses a dexterous grabbing method and system for a robot, and aims to solve the technical problems that an existing robot grabbing system is high in dependence on hand-eye calibration precision (especially translation vectors), and the grabbing capacity of objects without priori knowledge is limited. According to the method, a reference-transformation-correction control framework is provided, the pose of an unknown object is accurately recognized through a multi-stage visual segmentation strategy, and a preliminary grabbing target is generated based on the reference relation of one-time teaching; most importantly, the system starts closed-loop visual servo correction, and iteratively corrects the position by detecting the real-time pose of a tail end visual mark, so that translation errors in hand-eye calibration are eliminated, and high-precision alignment is realized; in addition, the system also integrates the functions of object posture adjustment and rapid rotation calibration. Through cooperative work of open-loop target generation and closed-loop position correction, dependence on hardware accurate calibration is effectively reduced, and the grabbing success rate and robustness are improved.
Owner:TIANJIN UNIV

Crop plot identification method and system based on time sequence vegetation characteristics

The invention relates to a crop plot identification method and system based on time sequence vegetation characteristics, and the method comprises the steps: constructing a normalized difference vegetation index time sequence according to a multi-temporal satellite remote sensing image, and generating a crop probability distribution diagram through a crop classification model; performing connected domain analysis on the crop probability distribution map, extracting the center of gravity of an effective connected domain as a forward attention point, scanning the crop probability distribution map by using a sliding window, and generating a reverse attention point; based on the forward attention point and the reverse attention point, segmenting the high-resolution remote sensing base map through a visual segmentation model, generating a candidate mask set, and screening to reserve a mask with the maximum area as a land parcel segmentation mask of the forward attention point; and combining all the plot segmentation masks to generate a crop plot identification graph, and converting the boundary of each plot segmentation mask into a geodetic coordinate sequence to obtain plot vector boundary data. The method improves the recognition accuracy, and achieves the automatic and precise segmentation of the boundary of the land parcel.
Owner:XIAN FEIFENG INTELLIGENT TECH CO LTD

Audio guidance visual segmentation method and device based on multi-granularity cross-modal coupling

The invention relates to the technical field of artificial intelligence and multimedia, in particular to an audio guidance visual segmentation method and device based on multi-granularity cross-modal coupling, and the method comprises the steps: extracting multi-level visual features and audio Mel spectrum features of a target video frame; performing intra-modal enhancement on the multi-level visual features and the audio Mel spectrum features to obtain enhanced multi-level visual features and enhanced audio Mel spectrum features; performing cross-modal fusion on the enhanced multi-level visual features and the enhanced audio Mel spectrum features to generate a semantic enhancement query vector; and training a pre-constructed Transform attention decoder by using the semantic enhancement query vector to generate a pixel-level segmentation mask, and fusing the pixel-level segmentation mask with the multi-level visual features to obtain a mask prediction result. Therefore, the problems of intra-modal noise interference, insufficient audio guidance, multi-source sound entanglement and the like in an existing audio-visual segmentation method are solved.
Owner:WUHAN UNIV

Control method, device and equipment of multi-mode large model robot and medium

The invention provides a multi-mode large model robot control method and device, equipment and a medium, and the method comprises the steps: carrying out the feature extraction, normalization processing and cross attention processing of a plurality of RGB images and depth images of a robot in a robot control model, and determining visual features; performing feature processing on the joint torque signal of the mechanical arm and the real-time current signal of the tail end gripper to determine a force feedback feature of the robot; feature coding processing is conducted on the task instruction of the robot, the current mechanical arm pose information and the current mechanical arm motion frequency, and task instruction features, visual segmentation features, pose features and frequency features are determined; and performing interactive attention processing and causal attention processing on the visual features, the force feedback features, the task instruction features, the visual segmentation features, the pose features and the frequency features in a robot control model, and outputting an action decision and visual information of the next moment. Therefore, the accuracy of action decision is improved.
Owner:SHENZHEN SHIHE ROBOTIC TECH CO LTD

Construction equipment situation analysis and event identification method and system based on visual segmentation

The invention relates to the technical field of construction safety monitoring, and particularly discloses a construction equipment situation analysis and event identification method and system based on visual segmentation. The method comprises the following steps: acquiring a construction equipment operation video stream through a roadside camera, and constructing an image and physical space mapping relation based on scene ground key feature points; a YOLOv11-Seg instance segmentation network is adopted to identify the contour of the equipment and a landing part, and a cross-frame motion track is tracked through a YOLOv11-Track algorithm; calculating a safe operation boundary by combining equipment geometric parameters, landing point coordinates and a space mapping relation; detecting personnel equipment invasion behaviors in real time, and implementing dynamic early warning according to kinematics parameters; and identifying operation event modes and anomalies through spatio-temporal feature analysis. According to the invention, high-precision perception of the operation situation of the construction equipment can be realized, the dynamic safety boundary model early warns the construction risk in real time, and spatial-temporal feature analysis provides a basis for intelligent management and efficiency improvement of the construction site.
Owner:GUANGZHOU UNIVERSITY

Industrial equipment operation analysis method and system based on machine vision

The invention relates to the technical field of machine vision, in particular to an industrial equipment operation analysis method and system based on machine vision. According to the method, the track information of the guide rail of the numerical control machine tool is obtained, and the three-dimensional coordinates and the timestamp sequence of the track points are established, so that the dynamic change of the track of the guide rail can be accurately described, the key track points are extracted through density analysis, and it is ensured that the perception of the track form change is more targeted. The time window division of the trajectory data is combined with the form change amplitude calculation, so that the quantification of the trajectory deviation trend is realized, and the trajectory anomaly can be quickly identified. In combination with stress analysis of a track deviation overrun area, an expansion path of fatigue damage can be determined, and a visual description of a damage propagation process is formed. Through extraction and analysis of a guide rail surface visual image, features of local cracks, wear and deformation can be matched with track anomalies, a visual segmentation boundary is further optimized, and it is ensured that details of structure changes are accurately extracted.
Owner:NANTONG INST OF TECH

Semi-supervised image segmentation method based on adaptive pixel subdivision

The invention provides a semi-supervised image segmentation method based on adaptive pixel subdivision, and relates to the technical field of computer vision and machine learning, and the method comprises the steps: firstly collecting image data according to a target scene, carrying out the preprocessing, dividing a training set and a test set, and marking part of data; preliminary segmentation is carried out by using a preset semantic segmentation model, a key pixel point region is screened through uncertainty analysis, and the segmentation precision is improved in combination with multi-scale features and iterative optimization. A teacher-student model framework is constructed, a teacher model is utilized to generate a pseudo tag, a high-confidence sample is screened to optimize a student model, and a contrast learning strategy is introduced to low-confidence pixels to enhance the generalization ability of the model. And finally, inputting the to-be-segmented image into the optimized student model, outputting a category matrix and generating a visual segmentation result, thereby improving the segmentation precision and boundary detail performance, and being suitable for a scene with limited annotation data.
Owner:INSPUR GENERSOFT CO LTD

Friction stir welding defect detection method fused with weak supervised learning

The invention provides a friction stir welding defect detection method fused with weak supervised learning, which belongs to the technical field of welding defect detection, and constructs a unified detection framework suitable for various weak labels such as points, frames, graffiti and the like by referring to the advantages of a large model SAM (Section Anything Model) in visual segmentation. The method comprises the following steps: firstly, converting different types of weak tags into an input format acceptable by SAM by adopting a prompt adapter module, and enhancing the prompt compatibility of a model; and secondly, abnormal region responses are eliminated through a response filter, and the recognition precision of the disguise or low-contrast defect region is improved in combination with a semantic matcher. A prompt self-adaptive knowledge distillation mechanism is further introduced, so that knowledge migration from the SAM model to the lightweight detection model is realized, and the feature learning ability of complex weld defects is enhanced. According to the method, high-quality defect detection can be realized in the weld defect image without a large number of accurate labels, and particularly, the method has remarkable advantages in the aspect of processing the welding defects with fuzzy boundaries, small sizes or weak contrast, and has wide industrial application value.
Owner:GUANGXI UNIV

A method and system for long distance rail identification and obstacle detection

The application discloses a long-distance rail recognition and obstacle detection method and system, a long-focus optical system is combined with a super-high-speed industrial camera on a hardware level to build a high-frame-rate and long-distance optical detection system, a region of interest for obstacle detection is demarcated by recognizing and visually segmenting left and right rail lines on a software level, and a target detection-based track obstacle intrusion detection algorithm and other advanced visual tasks are executed, and finally high-frame-rate recognition and detection of long-distance rails and track obstacles are realized. The Schmidt-Cassegrain optical system with a focal length of 2000mm is combined with a super-high-speed industrial camera with a resolution of 1920*1080 and 3000 frames per second to build a long-distance high-frame-rate optical detection platform and erect on a driving position of a train, railway conditions within a range of 6km are detected in real time, and serious accidents caused by a too long braking distance of the train due to high-speed driving can be prevented. Compared with traditional train driver eye observation and laser radar scanning imaging modes, a long-distance optical observation platform and a rear-end real-time detection system can realize a longer detection distance, better detection and recognition robustness and a higher frame rate.
Owner:BEIJING INST OF TECH

Visual segmentation large model data automatic labeling method combined with image prompt

The invention discloses a visual segmentation large model data automatic annotation method combined with image prompt, which comprises the following steps: extracting deep features of a prompt image and a to-be-annotated image by adopting a shared visual coding technology of deep learning, and carrying out cross-image feature matching and similarity calculation. And positioning a high-response area which is consistent with a prompt target in semantics in the to-be-labeled image, and automatically screening out an accurate prompt point set from the high-response area. Furthermore, based on the point set, the visual segmentation large model is guided to complete accurate segmentation and mask generation of the target, so that full-process automation from image prompting to final labeling is realized, and labeling efficiency and consistency are remarkably improved. In this way, the bottleneck problems that in the prior art, an automatic labeling method is insufficient in universality and the application of a segmentation large model must depend on human interaction are solved, and an effective technical approach is provided for large-scale, low-cost and high-quality visual segmentation data production.
Owner:HANGZHOU HUICUI INTELLIGENT TECH CO LTD

A bridge damage assessment method and system based on a visual segmentation model and a medium

This invention relates to the field of bridge structural safety inspection technology, and in particular to a bridge damage assessment method, system, and medium based on a visual segmentation model. The method includes: acquiring prior structural information and post-damage images of the bridge to be inspected; matching the image to be inspected generated based on the post-damage images with the prior structural information and a preset damage semantic knowledge base to generate multimodal prompts; inputting the fused multimodal prompts and the image to be inspected into a visual segmentation model with input parameters frozen, and adaptively iterating to obtain pixel-level segmentation masks for each damage candidate region; superimposing each pixel-level segmentation mask onto the image to be inspected, and inputting it along with the damage analysis prompt template into a visual language model to generate a structured damage assessment report. This invention utilizes the general segmentation capability of the visual segmentation model to achieve rapid and accurate pixel-level segmentation and recognition of sudden bridge damage, and automatically generates a structured damage assessment report.
Owner:TIANJIN ANXIN DIGITAL TECHNOLOGY CO LTD

Element identification method and device based on visual segmentation and storage medium

The invention relates to the technical field of image processing, in particular to an element recognition method and device based on visual segmentation and a storage medium, and the scheme comprises the steps: receiving a positioning parameter set for a to-be-recognized element; performing image acquisition on the to-be-identified element according to the positioning parameter set to obtain at least two to-be-identified images; performing feature recognition on the at least two to-be-recognized images to obtain first coordinate data of each to-be-recognized image in the at least two to-be-recognized images; performing coordinate conversion processing on the first coordinate data of each to-be-identified image to obtain converted second coordinate data; and obtaining target position information of the to-be-identified element according to the positioning parameter set and the second coordinate data. According to the method, through image acquisition, multi-coordinate system conversion and flexible parameter configuration, high-precision, high-adaptability and high-efficiency component identification is realized, the method is especially suitable for a mounting scene of a large component or a special-shaped component, and the reliability and economical efficiency of an automatic production line are remarkably improved.
Owner:SHENZHEN FAROAD INTELLIGENT EQUIP CO LTD

Cross-domain small sample wideband signal detection and identification method based on visual foundation large model

PendingCN122637056AVisual BasicAlgorithm
The application provides a cross-domain small sample wideband signal detection and recognition method based on a visual basic model, and relates to the cross technical field of electromagnetic signal processing and computer vision. The application reduces the interference of background noise in the wideband time-frequency graph on the candidate area by performing double condition screening based on target confidence and aspect ratio features on the preliminary candidate frame, and combining the field of view expansion in the frequency axis direction, and improves the incomplete coverage problem of the slender signal candidate frame. By extracting multiple intermediate layer features and the final layer feature map of the visual basic model for fusion in the channel dimension, the problem that the local texture information in the wideband time-frequency graph is smoothed or covered in the layer-by-layer abstraction process is solved. The parameters of the visual segmentation model and the visual basic model are kept frozen, and only the small sample fine tuning of the front region proposal network is performed, so that the calculation overhead caused by full fine tuning of the visual basic model during cross-domain small sample adaptation is avoided, and the method is suitable for deployment on devices with limited computing resources.
Owner:NORTHEASTERN UNIV CHINA

Full-autonomous ultrasonic robot real-time liver tumor tracking system

The invention provides a full-autonomous ultrasonic robot real-time liver tumor tracking system, and relates to the technical field of medical images and robots. According to the invention, through deep fusion of a multi-mode force touch sensing closed loop and an image segmentation visual closed loop, full-autonomous ultrasonic tracking of liver tumors is realized. Mechanical sensing ensures that the probe is always in stable contact with a patient in a proper posture and force, and visual segmentation feedback ensures that the probe is continuously aligned with and focuses on a tumor target. Under the synergistic effect of the two paths of feedback, the robot can automatically complete the whole process from probe placement, stable attachment, target recognition to real-time tracking, the uncertainty of manual operation is greatly relieved, and the safe, efficient and intelligent liver tumor real-time ultrasonic monitoring robot system is achieved.
Owner:SUZHOU INST OF BIOMEDICAL ENG & TECH CHINESE ACADEMY OF SCI

Structure multi-target deformation monitoring method based on visual segmentation and deblurring enhancement

The invention discloses a structure multi-target deformation monitoring method based on visual segmentation and deblurring enhancement, and relates to the technical field of engineering structure health monitoring. The method comprises the following steps: acquiring continuous time sequence multi-target images of different depth-of-field planes of a to-be-monitored structural plane; inputting an interactive prompt point into a pre-training segmentation model in the first frame to generate a multi-target mask, and removing a noise region based on connected domain analysis; cutting by adopting a self-adaptive multi-scale cutting strategy based on boundary judgment to obtain sub-graphs; performing image enhancement on the sub-image by adopting a pre-trained deblurring neural network; extracting a central point and an external rectangle coordinate of a target mask area in the sub-graph and uniformly mapping the central point and the external rectangle coordinate back to a full-graph coordinate system; and dynamically updating a cutting window of a subsequent frame based on the central point and boundary information of the first frame, and converting pixel-level displacement into physical unit or three-dimensional coordinate change by comparing the central point positions of the first frame and the subsequent frame and combining camera calibration parameters, thereby realizing high-precision continuous monitoring.
Owner:UNIV OF SCI & TECH BEIJING

Context compression method and system based on visual modality and medium

The invention provides a context compression method and system based on a visual modality and a medium, and the method comprises the steps: obtaining long text context data, and rendering the long text context data into a corresponding document image; performing context optical coding on the document image, and performing segmentation processing on the document image by adopting a visual segmentation model to obtain a plurality of initial visual tokens; performing down-sampling compression on the initial visual token based on a convolution compression algorithm to obtain a compressed token; capturing a long-distance dependency relationship among different compression tokens based on a global attention mechanism, extracting high-level semantic knowledge, and outputting a final visual token sequence; executing context optical decoding on the final visual token sequence, and fusing context optical decoding data with a text prompt input by a user to generate a target text; through visual modal compression, computing resources and memory occupation required for processing a long text are reduced, and relatively high text decoding precision can still be maintained while a high compression ratio is maintained.
Owner:SHENZHEN KUSI BIOTECHNOLOGY CO LTD

Visual segmentation counting method and system suitable for disorderly stacked parts

The invention relates to the field of image processing, and provides a visual segmentation counting method and system suitable for disorderly stacked parts, which can avoid the influence caused by a complex construction environment by carrying out dynamic Gamma correction and overlapping region enhancement processing on part images of multiple visual angles, highlight and enhance the characteristics of the overlapping region, and improve the accuracy of the visual segmentation counting of the disorderly stacked parts. Specific detail information of a shielding part and a small target is easier to recognize, recognition accuracy and robustness are improved, basic semantic segmentation is performed through a lightweight CNN model, recognition precision is guaranteed, meanwhile, calculation amount is greatly reduced, recognition efficiency is improved, and through dynamic threshold segmentation, seed point production and region growth, recognition efficiency is improved. According to the visual segmentation counting method for the disordered and stacked parts, the semantic and geometric features of the image are combined, the shielded and small-size parts are segmented more accurately, the omission ratio is effectively reduced, higher adaptability and detection precision are shown in a complex scene, and the accuracy and efficiency of the visual segmentation counting method for the disordered and stacked parts are improved.
Owner:JIANGXI XINCHUANGZHAN AUTOMOBILE & MOTORCYCLE PARTS CO LTD

A Scene Text Segmentation Method Based on an Improved SAM Visual Segmentation Model

This invention relates to the field of scene text segmentation, specifically a scene text segmentation method based on an improved SAM visual segmentation large model. Based on the SAM visual large model, this invention extracts text content perception features through an image content perception module and text edge perception features through a text edge perception module. Furthermore, the text feature fusion module extracts and calculates text edge perception feature maps, which are then added to the vectors requiring attention calculation before each self-attention calculation in the SAM encoder. This improves the accuracy of SAM in text segmentation and shortens the model training time while maintaining generalization.
Owner:ZHEJIANG UNIV OF TECH

Cross-station filtered and remaining area spray trajectory re-planning method and system

PendingCN122442672ACloud processingSimulation
The application discloses a kind of filtering and remaining area spray trajectory replanning method and system across station has been sprayed area, belong to industrial robot vision and automatic spraying technical field.This system is composed of mobile chassis, six-axis robot arm, terminal RGB-D camera, spray gun, station visual label and supporting function module;Robot generates standard spraying trajectory by visual segmentation, point cloud processing in first station and records sprayed area data, after moving across station, relies on visual label to complete multi-frame fusion relocation, calculates inter-station coordinate transformation matrix, maps historical sprayed area to current coordinate system and eliminates from real-time point cloud, re-plans serpentine spraying trajectory for remaining to be sprayed area, and accurately controls spray gun on-off in combination with trajectory segmentation logic.The application solves the problem that traditional mobile spraying robot is prone to repeated spraying, uneven paint film and repeated teaching when working across station, realizes multi-station continuous automatic spraying, and the trajectory runs smoothly with high spraying consistency.
Owner:ZHIYOUWUJIE (SHENZHEN) INTELLIGENT TECHNOLOGY CO LTD +2

Memory parking cross-layer identification method and device, electronic equipment and storage medium

The invention provides a memory parking cross-layer identification method and device, electronic equipment and a storage medium, and relates to the technical field of intelligent driving. In the method, a vehicle cross-layer state is determined through pitch angle features and a visual semantic segmentation result; the pitch angle features comprise a pitch angle mean value, a pitch angle variance and a difference value between the latest frame of pitch angle data and the oldest frame of pitch angle data in the sliding window, cross-layer identification failure caused by value abnormity due to direct use of the pitch angle data can be avoided, the correctness of the cross-layer state of the vehicle is ensured in combination with a visual segmentation result, and in addition, the accuracy of the cross-layer state of the vehicle is improved. Through a set of state switching logic, identification of the cross-floor state of the multi-floor underground parking lot (namely, the height of the current floor where the vehicle is located is determined according to switching information) is achieved, the accuracy is good, and the reliability is high.
Owner:NEUSOFT REACH AUTOMOBILE TECH (SHENYANG) CO LTD

A visual detection method and device for a ship segment large-plane laser derusting wall-climbing robot

This invention belongs to the field of shipbuilding and surface treatment technology, specifically relating to a visual inspection method and device for a laser rust removal climbing robot for large flat surfaces in ship sections. It is applied to secondary rust removal operations on the flat surfaces of ship sections during the manufacturing process, using a laser cleaning climbing robot. Particularly, it relates to a technical solution that uses visual rust detection results to identify rust conditions in real time and automatically set laser cleaning process parameters by matching them to a process parameter library. Furthermore, this invention also relates to a method for accurately determining the rust range using visual segmentation technology to plan the laser cleaning operation area, and a closed-loop control technology for real-time inspection of the rust removal effect using a vision system after laser cleaning. It also relates to lens protection measures for the camera module in the laser cleaning operation environment. This invention can be widely applied to automated laser cleaning operations of large flat structures in shipbuilding, marine engineering equipment, and steel structure surface treatment fields, improving the overall level of automation.
Owner:SHIPBUILDING TECHNOLOGY RESEARCH INSITITUTE (NO 11 INSTITUTE OF CSSC)

Electrolytic anode steel claw cleaning system

The invention discloses an electrolytic anode steel claw cleaning system. The electrolytic anode steel claw cleaning system comprises self-adaptive throwing chain cleaning, gas-solid coupling dust removal interlocking, anti-explosion visual segmentation scoring, optional robot fine cleaning, cross-station PLC interlocking release and NVR / digital twinning coaxial tracing. Based on load fingerprint and vision double judgment and pressure difference / air pressure trend guarding, the cleaning strength is adjusted in a self-adaptive mode, insufficient cleaning / excessive cleaning and secondary dust raising are restrained, and verifiable evidence export is supported.
Owner:QINGTONGXIA ALUMINUM GRP

A control method, device and equipment of a multi-modal large model robot and a medium

ActiveCN121290398BFeature extractionRgb image
The application provides a control method and device of a multimodal large model robot, equipment and a medium, comprising: performing feature extraction, normalization processing and cross attention processing on a plurality of RGB images and depth images of the robot in the robot control model to determine visual features; performing feature processing on the joint torque signal of the robot arm and the real-time current signal of the end gripper to determine the force feedback features of the robot; performing feature encoding processing on the task instruction, the current robot arm pose information and the current robot arm motion frequency to determine the task instruction features, the visual segmentation features, the pose features and the frequency features; and performing interactive attention processing and causal attention processing on the visual features, the force feedback features, the task instruction features, the visual segmentation features, the pose features and the frequency features in the robot control model to output the action decision and the visual information at the next moment. Thus, the accuracy of the action decision is improved.
Owner:SHENZHEN SHIHE ROBOTIC TECH CO LTD

Soil fertility determination method and system based on artificial intelligence vision

The present application belongs to the technical field of agricultural visual intelligent detection, and particularly relates to a soil fertility determination method and system based on artificial intelligence vision. The method comprises: obtaining information of a soil sample to be measured, and contacting soil extract with a color reaction tube of a fertility detection kit; collecting a target image of the kit through a rear camera of a mobile device, and collecting an ambient light reference image of a diffuse reflection reference piece through a front camera; positioning the color reaction tube, a blank control color block, a standard color block area and a positioning mark area through a visual segmentation network, and performing perspective correction; generating a three-reference chroma correction matrix according to an ambient light chroma vector, a standard color block measured chroma value and a blank area chroma value, and obtaining corrected chroma values of each detection unit; and matching the fertility determination value in combination with the batch number of the kit and the actual color development duration. The present application can reduce the influence of light, equipment and kit base differences on the interpretation result, and improve the consistency and traceability of the soil fertility determination result.
Owner:HUNAN CHEM VOCATIONAL TECH COLLEGE

Pipeline construction prefabricated part division method and system based on three-dimensional model

The invention relates to a pipeline construction prefabricated part division method and system based on a three-dimensional model, and the method comprises the following steps: constructing a pipeline construction three-dimensional model and a construction environment model, and displaying the prefabricated part division operation of a user on the three-dimensional model in real time through a visual division operation interface; on the basis of a preset pipeline element connection rule base, the physical connection relation between the pipeline elements is automatically checked in real time, and meanwhile the installation position interference state of the pipeline elements and the construction environment model is checked; performing analysis learning and training on historical project data based on a machine learning algorithm according to construction parameters and division operation set by a user, automatically generating a prefabricated part division scheme, and performing rationality check on the division scheme; and outputting integrated data containing detailed information of each prefabricated member. By constructing a three-dimensional model of pipeline construction and combining intelligent verification and a visual operation interface, intelligent division and optimal management of the pipeline prefabricated parts are achieved.
Owner:DMS CORP +1

Dynamic division method and system for multi-stage early warning collection area based on visual segmentation

The application discloses a dynamic division method and system for a multi-stage early warning winding area based on visual segmentation, and belongs to the technical field of machine vision, and aims to solve the defects of frequent calibration, weak anti-interference ability and lack of intelligent grading mechanism in traditional dynamic programming methods.The method comprises the following steps: adaptively triggering a camera to obtain an image through the method of taking frames in the air; performing real-time semantic segmentation on the current image based on a light-weight segmentation network model to obtain a time sequence mask of the current image as a current mask, the categories of the time sequence mask being two, namely, a coil and a winding wire; updating the mask based on time sequence perception, calculating the intersection and union ratio between the current mask and a historical reference mask, and triggering the update of a warning area model; when the update condition of the warning area is met, weighting and fusing the current mask and the historical reference mask to generate a final mask; and extracting key points from the final mask and grading to construct a winding warning area.
Owner:INSPUR QILU SOFTWARE IND

Methods and apparatus for bright-field cell image segmentation in fluorescence microscopy

This invention relates to a method and apparatus for bright-field cell image segmentation in fluorescence microscopy. Combining an improved two-dimensional OTSU threshold segmentation algorithm, it filters out noise points in single-cell images during image preprocessing and further refines cell segmentation using binary image mathematical morphology. Then, it segments adhered cells using a label-controlled watershed segmentation algorithm based on cell nucleus images. This invention improves segmentation results through image enhancement and refines them using various segmentation methods, effectively addressing issues such as weak edges, poor contrast, irregular cell shapes, and cell adhesion in bright-field cell images. Therefore, the overall visual segmentation effect is quite good.
Owner:HUAQIAO UNIVERSITY

A visual segmentation counting method and system suitable for jumbled stacked parts

The present application relates to the field of image processing, and proposes a visual segmentation and counting method and system suitable for disordered stacked parts, which avoids the influence of complex construction environment by performing dynamic Gamma correction and overlapping area enhancement processing on multiple view part images, simultaneously highlights and enhances the features of the overlapping area, makes it easier to identify the specific details of the occluded part and small target, improves the accuracy and robustness of identification, then performs basic semantic segmentation through a lightweight CNN model, greatly reduces the calculation amount while ensuring the identification accuracy, improves the identification efficiency, and further produces seed points and region growth through dynamic threshold segmentation, combines the semantic and geometric features of the image, more accurately segments the occluded and small-sized parts, effectively reduces the missed detection rate, and exhibits stronger adaptability and detection accuracy in complex scenes, and the present application improves the accuracy and efficiency of the visual segmentation and counting method of disordered stacked parts.
Owner:JIANGXI XINCHUANGZHAN AUTOMOBILE & MOTORCYCLE PARTS CO LTD

A remote sensing image water body extraction method and system based on texture topological feature guidance and visual segmentation large model cooperation

PendingCN122368785AGuaranteed puritySolve the problem of easily misidentifying high-brightness noise points as water bodiesThresholdingVisual perception
The application discloses a remote sensing image water body extraction method based on texture topological feature guidance and visual segmentation large model cooperation, comprising the following steps: acquiring a multi-modal remote sensing image, and calculating a statistical feature map; based on the statistical feature map and a preset brightness physical constraint threshold, a texture constraint mask and a brightness suppression mask are respectively constructed, and a spectral feature mask is selectively constructed, and then a potential water body mask is obtained based on two or three thereof; a target connected region is extracted from the potential water body mask, and a geometric extreme point in the target connected region is calculated based on a spatial topological analysis algorithm and is defined as a spatial position prompt; the multi-modal remote sensing image and the spatial position prompt are input into a visual segmentation large model which is pre-trained and activated, and a multi-level candidate mask and a corresponding prediction confidence score are output; and the mask with the highest score is selected as the final remote sensing image water body extraction result. Based on this, the application realizes the automatic extraction of the remote sensing image water body.
Owner:WUHAN UNIV

Multimodal pixel-level detection method, device, system, and storage medium for road surface cracks

This invention discloses a method, device, system, and storage medium for multimodal pixel-level detection of road cracks, comprising: constructing an automated data synthesis pipeline; constructing a multimodal road crack dataset; deeply fusing the general visual language model Qwen2.5-VL with a visual segmentation model through a cue injection mechanism to construct a multimodal pixel-level detection model for road cracks; designing a coordinate transformation module and a text projection module to encode and inject the explicit bounding box cues and implicit semantic cues generated by Qwen2.5-VL into the visual segmentation model; fine-tuning Qwen2.5-VL using a hierarchical symmetric low-rank adaptation strategy to enable it to learn professional knowledge in the crack detection field; and finally, outputting image-to-text understanding, visual location coordinates, and inference segmentation masks through the multimodal pixel-level detection model for road cracks. This invention addresses the shortcomings of visual detection methods lacking semantic understanding and large models lacking pixel-level detection accuracy.
Owner:CHENGDU UNIV OF INFORMATION TECH