Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

28 results about "Visual Basic" patented technology

Visual Basic is a third-generation event-driven programming language from Microsoft for its Component Object Model (COM) programming model first released in 1991 and declared legacy during 2008. Microsoft intended Visual Basic to be relatively easy to learn and use. Visual Basic was derived from BASIC and enables the rapid application development (RAD) of graphical user interface (GUI) applications, access to databases using Data Access Objects, Remote Data Objects, or ActiveX Data Objects, and creation of ActiveX controls and objects.

Road environment semantic segmentation method based on cross-modal difference modulation and visual basis model

The invention discloses a road environment semantic segmentation method based on cross-modal difference modulation and a visual basic model, and the method is based on a DINOv3 visual basic model of a frozen weight, and introduces a complementary difference fusion module and a bidirectional context flow alignment module through the design of double-flow frozen coding and difference perception interaction. According to the method, a cross attention weight is generated by calculating a local high-frequency difference, so that transverse anti-noise complementation of RGB texture information and a Depth geometric structure is realized, and single-mode noise pollution is effectively inhibited; the two-way cross-scale alignment mechanism of top-down space mask guidance and bottom-up channel feature feedback is constructed by the two-way cross-scale alignment mechanism, and the scale barrier of deep and shallow layer features is broken. Therefore, joint feature representation with deep synergy of semantics and structures is obtained, inhibition of the model to environmental noise and reservation of tiny object details are balanced in a complex road scene, and pixel-level semantic segmentation with high generalization and high precision is realized.
Owner:GUANGDONG UNIV OF TECH

Sparse view angle pose-free scene reconstruction method and system based on 3DGS

The invention relates to a sparse view angle pose-free scene reconstruction method and system based on 3DGS, belongs to the field of computer vision and three-dimensional reconstruction, and solves the problems of easy failure and poor geometric consistency in a sparse view angle or weak texture environment due to dependence on accurate camera pose priori. Constructing a double-flow sensing module containing semantic flow and geometric flow, extracting semantic features by using a visual basic model, and extracting an explicit geometric corresponding relation by using a dense feature matching network, so as to regress relative camera pose under pose-free priori; a geometric guidance depth refinement module combining a potential diffusion model architecture and Pluecker ray coding is introduced, and scale fuzziness of monocular depth estimation is eliminated through a depth residual prediction mechanism; and based on the micronizable Gaussian rasterization, performing end-to-end optimization by using a self-supervised loss function including rendering consistency, reprojection and epipolar geometric constraint. According to the method, high-fidelity three-dimensional reconstruction is realized without supervision of external parameter true values, and geometric stability and rendering quality in a complex scene are improved.
Owner:BEIJING UNIV OF TECH

Multi-modal visual position identification reordering method and system based on guidance

The invention relates to the technical field of visual position recognition, and particularly discloses a multi-modal visual position recognition reordering method and system based on guidance, and the method comprises the steps: obtaining a query image, and retrieving a plurality of candidate images based on a pre-trained visual basic model and the query image; constructing a composite multi-modal prompt object, wherein the composite multi-modal prompt object comprises an image pair formed by the query image and the current candidate image, and an instruction text used for guiding a multi-modal large language model to perform visual comparison; outputting a structured similarity judgment result, wherein the result comprises a quantitative similarity score; and sorting based on the similarity scores corresponding to all the candidate images, and determining the candidate image with the highest score as an optimal matching result. Through combination of guiding type prompt engineering and structured output, an intermediate text generation link is avoided fundamentally, and the calculation efficiency is improved while the fidelity of all original visual information is reserved.
Owner:SHENZHEN 1024 ROBOT TECHNOLOGY CO LTD

Monocular 3D object detection method for realizing depth enhancement based on visual basic model, electronic equipment and readable storage medium

The invention belongs to the technical field of computer vision, and particularly discloses a monocular 3D object detection method for realizing depth enhancement based on a visual basic model, electronic equipment and a readable storage medium, and the method comprises the steps: S1, building a data set: employing a monocular camera to collect a pavement scene, and obtaining an RGB image in the pavement scene; s2, image preprocessing: preprocessing the RGB image for subsequent feature extraction and depth estimation; s3, performing feature extraction by adopting a dual-backbone network: performing visual semantic feature extraction on the preprocessed RGB image by using DINOv2; performing depth feature extraction on the preprocessed RGB image by using a DPT head; s4, generation of depth perception query points: inputting the visual semantic features and the depth features into a DETR network to generate the depth perception query points; and S5, target detection output: using an MLP-based detection head to obtain information of the category, the size, the center point position, the depth, the 3D size and the direction of the object.
Owner:SHANGHAI UNIV

Industrial surface defect detection method fusing large model and domain knowledge

The invention provides an industrial surface defect detection method fusing a large model and domain knowledge, and relates to the technical field of industrial detection, and the method specifically comprises the steps: designing a double-branch knowledge injection network model which comprises a backbone network, a side branch network, a texture encoder, a curvature extraction network and a detection head network, extracting general visual features by using a visual large model in the backbone network, then adding and fusing the general visual features with supplementary features extracted by the bypass network to obtain deep features, extracting texture features by using a texture encoder, and embedding texture domain knowledge into the deep features and the texture features in a cross attention mode to obtain a deep texture domain; optimizing a loss function according to the stress attention map; positioning and classifying defects by using a detection head network; and training and evaluating the double-branch knowledge injection network model. According to the technical scheme, the problems that in the prior art, the detection performance of a visual basic model in an industrial scene is reduced due to field differences, and the small sample generalization ability is weak are solved.
Owner:SHANDONG UNIV OF SCI & TECH

Visual intelligent monitoring method and system for operation of belt conveyor

The invention relates to a visual intelligent monitoring method and system for operation of a belt conveyor. The method comprises the following steps: collecting a video stream through visual sensing equipment deployed along a belt conveyor; performing dynamic image quality evaluation and enhancement at an edge computing node, and operating a lightweight deep learning model to detect typical anomalies in real time; meanwhile, uploading the data to a cloud end, extracting high-dimensional features by using a visual basic model, and comparing the high-dimensional features with a dynamic feature distribution model to identify unknown anomalies; when an exception is found, data from the non-visual sensing module in the same time period are automatically called and analyzed to perform collaborative verification; and through a verification result, carrying out fusion decision making and generating final diagnosis and early warning information. The system comprises a visual perception module, an edge computing gateway, a non-visual sensing module and a cloud intelligent analysis platform. According to the invention, omnibearing and high-reliability intelligent monitoring and early warning of the running state of the belt conveyor are realized.
Owner:CCCC MECHANICAL & ELECTRICAL ENG

Small sample PCBA component target detection method and system based on visual basic model

The invention belongs to the field of PCBA quality detection, and relates to a small sample PCBA component target detection method and system based on a visual basic model. The method comprises the following steps: acquiring a pre-positioning prompt box from a PCBA image to be detected by using a positioning basic model, and acquiring a single-scale freezing feature of the image by using a classification basic model; expanding the single-scale features into multi-scale features for pre-positioning and classification tasks; pre-positioning prompts are screened and filtered, and query point features are obtained by using a feature alignment query sampling method; and predicting a classification result and a regression result for each query point feature according to the multi-scale positioning feature, the classification feature, the query point feature and the prompt box. According to the method, the high-precision model is quickly and finely adjusted under the condition that training data is scarce, the training cost is reduced, and the requirement for quick switching of production lines in small-batch PCBA production is effectively met.
Owner:HUAZHONG UNIV OF SCI & TECH

A method and system for zero-shot image segmentation based on object three-dimensional models

PendingCN122157261AHigh precisionEfficient automated segmentationBiological modelsKnowledge based modelsVisual BasicImage segmentation
The application relates to the field of image segmentation technology in computer vision, in particular to a zero-shot image segmentation method and system based on a three-dimensional model of an object. The method comprises the following steps: for any untrained object, a three-dimensional model of the object is rendered into a two-dimensional RGB reference image under a plurality of preset discrete viewing angles by using a graphics rendering engine, the reference image is input into a visual basic model DINOv3, a global feature vector representing semantic information of the object is extracted, and a reference feature library is constructed; a target image to be segmented is input into a visual basic model SAM2, and a plurality of candidate object masks are generated; for each candidate mask, an object image corresponding to the candidate mask is cropped from the target image, and the object image is input into the DINOv3 to extract a semantic feature vector of the candidate object; by calculating the cosine similarity between the candidate object feature vector and each template feature vector in the reference feature library, the top similar degrees with the highest values are selected, and an arithmetic mean value of the similar degrees is calculated, and the mean value is taken as the classification confidence of the candidate object mask; the class label of each candidate object mask is determined according to the confidence, and the class label of the target object and a corresponding pixel-level segmentation mask are output.
Owner:HANGZHOU HUXIYUN BAISHENG TECH CO LTD

A method for generating a sentiment data graph based on natural language understanding

ActiveCN121706786BSemantic analysisBiological modelsGraphicsVisual Basic
This invention discloses a method for generating emotional data graphics based on natural language understanding, comprising the following steps: collecting emotional text data, brand visual basic element data, and cross-cultural context description data, and processing the collected data; performing natural language understanding on the processed emotional text data; constructing a multi-domain visual semantic tensor flow that evolves over time; constructing a cross-cultural intent field and binding the parameters of the cross-cultural intent field to the multi-domain visual semantic tensor flow; generating a preliminary emotional graphic structure sequence through a generative visual network; dynamically adjusting the preliminary emotional graphic structure sequence; and outputting it to different media scenarios to obtain emotional data graphics. This invention integrates natural language understanding and multimodal generation technologies to achieve emotion-semantic driven graphics generation, possessing the advantages of semantic accuracy, cultural adaptation, and visual coherence.
Owner:DALIAN POLYTECHNIC UNIVERSITY

Cross-domain small sample semantic segmentation method and device, equipment and medium

The invention discloses a cross-domain small sample semantic segmentation method and device, equipment and a medium. According to the cross-domain small sample semantic segmentation method, a hierarchical data set is constructed through a controllable style offset synthesis image, and a style feature extraction and weight generation network is trained; a dynamic weight is generated based on the difference between the target domain image style representation and the source domain reference, feature layer level modulation is performed on the visual basis model encoder features, and the feature alignment problem is accurately solved; image semantic information is fused and supported through a memory attention mechanism, so that the problem of small sample semantic sparsity is effectively relieved; meanwhile, the synthetic data accurately controls style offset through a stable diffusion model, and uncontrollability of traditional generative data expansion is avoided. According to the method, the cross-domain small sample semantic segmentation performance is remarkably improved, and the problems of prompt mismatching, feature dislocation and out-of-control data expansion caused by domain offset are solved.
Owner:SHENZHEN UNIV

Fine adjustment method, system and equipment for visual basic model and storage medium

The invention discloses a fine tuning method, a fine tuning system and fine tuning equipment for a visual basic model, and a storage medium, which are corresponding schemes, and the related scheme aims to solve the problems of overlarge GPU video memory consumption, high quantitative perception fine tuning overhead and the like in the fine tuning and deployment process under the condition of limited resources in the existing parameter efficient fine tuning scheme. According to the scheme, through a sub-network adapter quantization perception fine tuning and block-level activation quantizer fine tuning strategy, video memory consumption in a fine tuning stage is remarkably reduced, and meanwhile, the calculation complexity consistent with that of an original backbone network in a reasoning stage is kept; therefore, a feasible and efficient solution is provided for low-resource deployment of the visual basic model in diversified downstream tasks.
Owner:UNIV OF SCI & TECH OF CHINA

Complex structure wheel forging rolling whole process digital design and performance prediction method

The invention discloses a digital design and performance prediction method for the whole forging and rolling process of a wheel with a complex structure, and belongs to the technical field of metal plastic forming and digital simulation. In order to overcome the defects of an existing method, an integrated forming module is designed; the invention provides a parametric modeling method, based on Visual Basic and SolidWorks secondary development, wheel parameters in an Access database are accessed, and three-dimensional modeling and assembling of all molds and blanks are automatically completed; a structure performance prediction method is provided, a parameterized model is automatically imported into finite element software, a microstructure model and a cellular automaton are integrated through secondary development, macro-micro coupling calculation is carried out, and microstructure distribution after the wheel is formed is predicted. According to the method, full-process digitization and automation from wheel design parameter input to final structure performance prediction are achieved, the research and development efficiency and the forming quality prediction precision of the complex-structure wheel are remarkably improved, and a core technical support is provided for optimal manufacturing of rail transit key components.
Owner:TAIYUAN UNIVERSITY OF SCIENCE AND TECHNOLOGY

Visual question-answering method and device based on multiple modes, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to the fields of financial science and technology and medical health, and discloses a multi-modality-based visual question and answer method, device, equipment and medium, and the method comprises the steps: obtaining an input image and a question, and inputting the image and the question into a visual language coding model for coding to generate a multi-modality feature; inputting the multi-modal features into an answer reasoning module to generate answers and deep reasoning features corresponding to the answers; inputting the multi-modal features into a basic principle generation module to generate guide features, and inputting the guide features into a large language model to generate a text basic principle; extracting features of the text basic principle to obtain context features, and generating a visual basic principle through an object detector according to the context features, the image and the depth reasoning features; and taking the visual basic principle, the text basic principle and the answers as final interpretable answers. And the transparency and the credibility of the visual question-answering system are effectively improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

RGB-D video salient target detection method based on depth-guided adaptive query

The invention relates to an RGB-D video salient target detection method based on depth guidance adaptive query, which comprises the following steps: firstly, constructing a parallel adapter structure based on depth guidance, and adopting a parallel jump structure to remarkably reduce the video memory overhead brought by gradient return during fine tuning, so that an SAM can automatically segment a salient target without any prompt; and meanwhile, multi-modal fusion of RGB-D features is realized by taking the depth map as guidance. And secondly, providing a query-driven time sequence memory module, replacing a huge Memory Bank in the SAM2 with frame-level query and video-level query, and realizing lightweight time sequence modeling. By introducing and improving a visual basic model structure, efficient and accurate video salient target segmentation is realized under the condition of no artificial prompt, and meanwhile, the video memory consumption and the calculation complexity of the model are remarkably reduced.
Owner:HANGZHOU DIANZI UNIV

Virtual object special effect logic editing method and device, equipment and medium

This invention discloses a method, apparatus, device, and medium for editing virtual object special effects logic. The method includes: receiving a configuration file and parsing it to obtain a configuration database; determining a selected object corresponding to selected information in an imported file; performing logical configuration on the selected object according to the configuration database to generate visual basic logical information; adjusting the parameters of the basic logical information according to parameter modification information; obtaining candidate special effects actions for the selected object from the configuration database; configuring special effects actions on the adjusted basic logical information according to action selection information; editing the configuration information in the configuration database; and generating a configuration editing file. This method can parse the configuration file to obtain a configuration database and perform logical configuration on the selected object, generating visual basic logical information for efficient parameter adjustment and special effects action configuration, thus improving the efficiency and accuracy of editing the characteristic logical information in the configuration file.
Owner:HANGZHOU WIZARD GAME TECH CO LTD

Machine learning model input monitor

The invention relates to a method, in particular a computer-implemented method, for monitoring the performance of a trained machine learning model, a corresponding monitoring system, a computer program and a computer-readable medium. The method comprises the following steps: providing a trained visual basic model; providing a training data set for training the machine learning model; for each training image in the training data set, using the visual basic model to determine a training data feature vector, and determining the distribution of the training data feature vector; receiving an image depicting a scene; determining an image feature vector of the image by using the visual basic model; calculating a log-likelihood value of the distribution of the image feature vector to the training data feature vector; if the image differs from the distribution of the training data feature vectors based on the analysis of the log-likelihood values, an alert is provided.
Owner:CONTINENTAL AUTONOMOUS MOBILITY GERMANY GMBH

A frequency-space dual-domain joint fine-tuning visual base model network method for remote sensing image domain generalization semantic segmentation

The application discloses a kind of frequency-space dual-domain joint fine-tuning visual basic model network methods for remote sensing image domain generalization semantic segmentation, which enhances the consistency of cross-domain similar features by adaptive frequency selection in the middle features of the frequency domain fine-tuned model, and improves the discriminant ability of the feature cluster boundary in the deep features of the spatial domain fine-tuned model to achieve robust domain generalization semantic segmentation performance. Through four experimental settings of three public datasets Potsdam, Vaihingen and LoveDA, the mIoU index of the present method is improved by 1.64% and 1.09% on average compared with the most advanced method, and the model parameter amount is only 4.22M, which balances the accuracy and computational efficiency.
Owner:BEIJING INST OF TECH

A zero-shot 3D model recognition method

PendingCN122347799APattern recognitionVisual Basic
The application discloses a zero sample three-dimensional model recognition method, comprising the following steps: constructing an auxiliary information embedded feature enhancement module, extracting text features and image features by using a visual basic model, and extracting point cloud features by using a learnable point cloud encoder; embedding the angle and attribute information of the image into the extracted image features as auxiliary information to obtain final image enhanced features; constructing a text guided feature adaptive selection module, calculating the average similarity score of the image enhanced features and the corresponding text features, and dynamically updating by using an exponential moving average strategy to obtain a global similarity score; adaptively selecting the image features by using the global similarity score; designing a cross-modal feature alignment mechanism constrained by a joint contrast loss, calculating the cross-modal contrast loss among the text, image and point cloud features; and realizing zero sample recognition of the three-dimensional model by measuring the distance between the three-dimensional model point cloud features and the text features of the category candidates.
Owner:TIANJIN UNIV

Task universal visual model construction method based on self-supervised representation

According to the task universal visual model construction method based on self-supervised representation, the potential feature space of a self-supervised visual basic model is used as the basis of unified representation, the reconstruction quality is improved by introducing a lightweight residual module, and a diffusion model is combined as a generation backbone, so that the task universal visual model construction method based on self-supervised representation is realized. Therefore, perception and generation tasks are simultaneously supported in a unified potential space. Compared with an existing method, the method has the advantages that the training and storage cost is remarkably reduced while the generation quality is guaranteed, higher mobility and expansibility are achieved, and the practical application requirements of various visual tasks can be efficiently met.
Owner:TSINGHUA UNIVERSITY

An industrial surface defect detection method fusing large models and domain knowledge

The application provides an industrial surface defect detection method fusing a large model and field knowledge, relates to the technical field of industrial detection, and specifically comprises the following steps: designing a double-branch knowledge injection network model, including a backbone network, a side branch network, a texture encoder, a curvature extraction network and a detection head network; extracting general visual features by using a visual large model in the backbone network, then adding and fusing the general visual features and supplementary features extracted by the side branch network to obtain deep features; extracting texture features by using the texture encoder; embedding texture field knowledge by using the deep features and the texture features in a cross attention manner; and optimizing a loss function according to a stress attention graph; positioning and classifying defects by using the detection head network; and training and evaluating the double-branch knowledge injection network model. The technical scheme of the application overcomes the problems of the existing technology, such as the decline of detection performance in an industrial scene caused by field differences and the weak small sample generalization ability of a visual basic model.
Owner:SHANDONG UNIV OF SCI & TECH

Method and system for detecting boundary enhancement based on visual base model feature adaptation

The application belongs to the field of remote sensing change detection, and discloses a detection boundary enhancement method and system based on visual basic model feature adaptation. The application performs adaptive processing on the features extracted from the visual basic model, so that the features are converted from general semantic representation to remote sensing change detection task sensitive representation. While maintaining strong semantic discrimination ability, the application effectively suppresses non-change interference such as illumination difference, seasonal change, sensor imaging difference, etc., improves the stability of change area determination from the source, and reduces false changes and missed detection phenomena. The application introduces a double-time multi-scale feature pyramid structure, so that large-scale features are responsible for the overall consistency determination of the change area, and small-scale features focus on small structures and complex edges, and the change area determination is stable and fine in space level.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Green data center facade facing material aided design system and method

The invention relates to a green data center facade facing material aided design system and method. The system is developed based on a Visual Basic programming language, adopts a Windows application program architecture, and comprises a material data management module used for collecting and sorting basic data of exterior facing materials and providing dynamic classification and keyword retrieval functions of the materials; the material combination and cost calculation module is used for obtaining comprehensive cost information by combining material market price information according to the material combination configured by the user, the proportion of each material and the selection condition of the thermal insulation material; the visual display module is used for configuring a high-definition picture and detailed attribute description for each material so that a user can click the picture to check detailed attribute information of the material, and providing three-dimensional effect display of the material; and the design auxiliary module is used for providing a rapid comparison function of design schemes, and comparing the cost and effect of different schemes based on the real-time adjustment of material combination and proportion or thermal insulation material selection of a user.
Owner:CHINA MOBILE GROUP DESIGN INST +1

A small sample detection method based on a visual base model

The application discloses a small sample detection method based on a visual basic model and relates to the technical field of computer vision. The method comprises the following steps: extracting a query picture visual feature map and obtaining foreground and background prototype vectors by using a pre-trained visual basic model; mapping the prototype vectors to the visual feature map to generate a category and background projection response map; obtaining a spatial weight by splicing and normalization based on the projection response map and generating a conditional feature map by transformation; enhancing the feature by using an inflation grouping convolution module and generating a prediction bounding box with confidence; determining a candidate box set by a dynamic selection strategy; constructing a multi-scale support region and extracting local and context features in an inference stage; and determining a final class label of the candidate box according to the highest frequency category in the classification results of each support region after feature interaction. The application aims to overcome the dependence on a basic class pre-trained detector and complex background interference, realize candidate box generation and robust prediction, and thus improve the precision and stability of small sample detection.
Owner:HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES

Three-dimensional scene feature embedding generation method based on 3DGS and multi-model memory

The invention provides a three-dimensional scene feature embedding generation method based on 3DGS and multi-model memory, and the method comprises the steps: firstly obtaining a to-be-processed multi-view two-dimensional image and a camera pose parameter thereof, obtaining a sparse three-dimensional point cloud, and then initializing an anchor point and a Gaussian ball; utilizing different visual basic models to extract high-dimensional features of the multi-view two-dimensional image to construct a feature memory bank; performing parameter compression, and performing 3D Gaussian splash rendering to obtain an RGB image and a main query vector of a current rendering view angle; and fusing the main query vector with the high-dimensional features in the feature memory library by using a Gaussian attention mechanism, and outputting three-dimensional scene feature embedding aligned with each visual basic model. The multi-modal feature embedding generation framework provided by the invention has good expandability and compatibility, can be adapted to multiple different types of pre-training basic models, and supports collaborative integration of multi-source modal features such as vision, language, depth and the like, so that unified three-dimensional scene semantic representation is realized.
Owner:HANGZHOU DIANZI UNIV

A remote sensing image interpretation method based on a lightweight visual base model

The application provides a remote sensing image interpretation method based on a lightweight visual basic model, a double-time-phase remote sensing image change detection model Edge-CD constructed by the method comprises an EdgeSAM encoder and a boundary reinforcement decoding head; the EdgeSAM encoder adopts an EdgeSAM image encoder as a feature extraction backbone, and a space-time feature alignment module is embedded in the EdgeSAM image encoder, so that the perception ability of the model to real changes is effectively enhanced, and the common false change problem is relieved. A boundary enhancement module is designed in the boundary reinforcement decoding head, so that the fitting degree of the boundary contour of the detection result to the real change is improved. In addition, in view of the problem that the traditional loss function is insufficient for optimizing the edge region, a boundary perception loss function is designed, a higher weight is given to the neighborhood region of the real change boundary, the gradient response of the model in the boundary region is enhanced, the edge false detection and the edge missing detection are effectively relieved, and the segmentation precision and the geometric integrity of the change boundary are further improved.
Owner:HUNAN INST OF WATER RESOURCES & HYDROPOWER RES +1

Weak texture insulator discharge trace identification method based on large model driving

The invention discloses a weak texture insulator discharge trace identification method based on large model driving, and the method comprises the steps: carrying out the multi-angle and multi-illumination-condition directional shooting of an insulator through an unmanned plane, and obtaining an insulator image set; and performing spatial domain and frequency domain feature enhancement on each image in the image set to generate a multi-channel feature map. An improved visual basic model is constructed, a task migration plug-in is combined, a multi-channel feature map is processed, a discharge trace segmentation mask is generated, the Monte Carlo Dropout technology is adopted to calculate and predict uncertainty, a low-confidence insulator image is automatically identified, and manual verification is carried out. A verification result is fed back to the task migration plug-in for incremental training, and the model performance is optimized. And completing automatic identification of insulator discharge traces by using the trained improved visual basic model. According to the method, the recognition precision of the weak texture discharge trace can be remarkably improved, manual intervention is reduced, and the method has relatively high engineering practicability.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Vision-based model oriented domain adaptation adversarial attack method and system

The present application belongs to the technical field of computer vision and attack and defense resistance, and provides a domain adaptation attack method and system for a visual basic model, and proposes an attack sample generation method capable of migrating attacks on unknown fine-tuning models by using an open-source basic model under a black-box condition. The attack process consists of two stages of offline domain alignment initialization and online instance-level adaptation. In the offline stage, the basic model is used to mine general vulnerabilities on a proxy data set to generate domain-aligned general perturbations; in the online stage, the general perturbations are introduced into the optimization process through momentum hot start, the structural gradient alignment module is used to guide the attack to focus on significant boundaries, and the amplitude deviation simulation module is combined to simulate the feature drift caused by fine-tuning in the frequency domain. The problems of domain offset and parameter difference between the open-source basic model and the specific domain model after fine-tuning are solved.
Owner:SHANDONG UNIV

Low-altitude remote sensing image small target interpretation and iterative correction method, equipment and medium

The invention discloses a low-altitude remote sensing image small target interpretation and iterative correction method and device and a medium, and the method comprises the steps: employing a pre-trained core algorithm model, employing a sliding window strategy to carry out the block reasoning prediction of an ultra-large-resolution low-altitude remote sensing image, and carrying out the vectorization of a result; with the help of GeoServer or GIS service, a prediction image combining image slices and vectorization prediction results is issued at an application front end; manually correcting a low-confidence or suspected error target in the predicted image, and writing a correction result into a vector database; and bringing the corrected prediction image into a training set, and carrying out incremental or fine training on the core algorithm model. According to the method, through fusing DINOv3 self-supervised visual basic model, improved ViT-Adapter and Cascade-RCNN multi-stage detection, high-precision automatic identification and continuous optimization of sparse small targets in an ultra-large remote sensing image are realized, and the method is especially suitable for low-altitude remote sensing image processing scenes with high resolution, large file volume, complex ground features and sparse targets.
Owner:ZHONGKE XINGTU DIGITAL EARTH HEFEI CO LTD