Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

49 results about "Visual Basic" patented technology

Visual Basic is a third-generation event-driven programming language from Microsoft for its Component Object Model (COM) programming model first released in 1991 and declared legacy during 2008. Microsoft intended Visual Basic to be relatively easy to learn and use. Visual Basic was derived from BASIC and enables the rapid application development (RAD) of graphical user interface (GUI) applications, access to databases using Data Access Objects, Remote Data Objects, or ActiveX Data Objects, and creation of ActiveX controls and objects.

Vision-language interaction-based remote sensing image open vocabulary segmentation method and system

The invention discloses a remote sensing image open vocabulary segmentation method and system based on vision-language interaction. According to the method, the advantages of a multi-modal large language model and a semantic segmentation network based on a visual basic model are cooperatively utilized, language-pixel two-way mapping is taken as a core, five types of marks of images, texts, categories, objects and segmentation are introduced as carriers of cross-modal information, and three types of cross-modal fine-grained information interaction modules are taken as bridges, so that the cross-modal information interaction is realized. Bidirectional mapping and alignment of fine-grained information of the multi-modal large language model and the semantic segmentation network are realized, the open vocabulary segmentation capability of the semantic segmentation model is improved, and the method can adapt to different remote sensing scenes and category definitions. The method has the following advantages: the method has high performance, strong generalization ability and good expansibility, can provide any category of semantic segmentation maps for unlabeled target domain images based on instructions, and has high application value in the aspects of urban planning, map making, disaster response and the like.
Owner:WUHAN UNIV

Surgical instrument segmentation method based on double prior guidance networks

The invention discloses a surgical instrument segmentation method based on a double prior guidance network, and the method specifically comprises the steps: constructing a hybrid encoder composed of a visual basic model and a state space model, and carrying out the coding processing of an input surgical endoscope image; the visual basic model is an SAM2-Hira model; the state space model is a MambaVision model; an adapter layer composed of a standard adapter and a Mama enhancement adapter is arranged between the coding layer and the decoding layer, and coding features are processed; setting a multi-context guide decoder, and performing multi-scale feature recovery and mask segmentation; a hybrid encoder, a Mama enhancement adapter and a multi-context guidance decoder form a double-priori guidance network DPG-Net, and high-precision surgical instrument segmentation is realized by using the DPG-Net. Accurate visual support is provided for the surgical robot, and the accuracy and safety of surgical operation are greatly improved.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

Road environment semantic segmentation method based on cross-modal difference modulation and visual basis model

The invention discloses a road environment semantic segmentation method based on cross-modal difference modulation and a visual basic model, and the method is based on a DINOv3 visual basic model of a frozen weight, and introduces a complementary difference fusion module and a bidirectional context flow alignment module through the design of double-flow frozen coding and difference perception interaction. According to the method, a cross attention weight is generated by calculating a local high-frequency difference, so that transverse anti-noise complementation of RGB texture information and a Depth geometric structure is realized, and single-mode noise pollution is effectively inhibited; the two-way cross-scale alignment mechanism of top-down space mask guidance and bottom-up channel feature feedback is constructed by the two-way cross-scale alignment mechanism, and the scale barrier of deep and shallow layer features is broken. Therefore, joint feature representation with deep synergy of semantics and structures is obtained, inhibition of the model to environmental noise and reservation of tiny object details are balanced in a complex road scene, and pixel-level semantic segmentation with high generalization and high precision is realized.
Owner:GUANGDONG UNIV OF TECH

Sparse view angle pose-free scene reconstruction method and system based on 3DGS

The invention relates to a sparse view angle pose-free scene reconstruction method and system based on 3DGS, belongs to the field of computer vision and three-dimensional reconstruction, and solves the problems of easy failure and poor geometric consistency in a sparse view angle or weak texture environment due to dependence on accurate camera pose priori. Constructing a double-flow sensing module containing semantic flow and geometric flow, extracting semantic features by using a visual basic model, and extracting an explicit geometric corresponding relation by using a dense feature matching network, so as to regress relative camera pose under pose-free priori; a geometric guidance depth refinement module combining a potential diffusion model architecture and Pluecker ray coding is introduced, and scale fuzziness of monocular depth estimation is eliminated through a depth residual prediction mechanism; and based on the micronizable Gaussian rasterization, performing end-to-end optimization by using a self-supervised loss function including rendering consistency, reprojection and epipolar geometric constraint. According to the method, high-fidelity three-dimensional reconstruction is realized without supervision of external parameter true values, and geometric stability and rendering quality in a complex scene are improved.
Owner:BEIJING UNIV OF TECH

Multi-modal visual position identification reordering method and system based on guidance

The invention relates to the technical field of visual position recognition, and particularly discloses a multi-modal visual position recognition reordering method and system based on guidance, and the method comprises the steps: obtaining a query image, and retrieving a plurality of candidate images based on a pre-trained visual basic model and the query image; constructing a composite multi-modal prompt object, wherein the composite multi-modal prompt object comprises an image pair formed by the query image and the current candidate image, and an instruction text used for guiding a multi-modal large language model to perform visual comparison; outputting a structured similarity judgment result, wherein the result comprises a quantitative similarity score; and sorting based on the similarity scores corresponding to all the candidate images, and determining the candidate image with the highest score as an optimal matching result. Through combination of guiding type prompt engineering and structured output, an intermediate text generation link is avoided fundamentally, and the calculation efficiency is improved while the fidelity of all original visual information is reserved.
Owner:SHENZHEN 1024 ROBOT TECHNOLOGY CO LTD

Monocular 3D object detection method for realizing depth enhancement based on visual basic model, electronic equipment and readable storage medium

The invention belongs to the technical field of computer vision, and particularly discloses a monocular 3D object detection method for realizing depth enhancement based on a visual basic model, electronic equipment and a readable storage medium, and the method comprises the steps: S1, building a data set: employing a monocular camera to collect a pavement scene, and obtaining an RGB image in the pavement scene; s2, image preprocessing: preprocessing the RGB image for subsequent feature extraction and depth estimation; s3, performing feature extraction by adopting a dual-backbone network: performing visual semantic feature extraction on the preprocessed RGB image by using DINOv2; performing depth feature extraction on the preprocessed RGB image by using a DPT head; s4, generation of depth perception query points: inputting the visual semantic features and the depth features into a DETR network to generate the depth perception query points; and S5, target detection output: using an MLP-based detection head to obtain information of the category, the size, the center point position, the depth, the 3D size and the direction of the object.
Owner:SHANGHAI UNIV

Handheld object three-dimensional reconstruction method based on three-dimensional generation prior and semantic consistency

The invention discloses a handheld object three-dimensional reconstruction method based on three-dimensional generation prior and semantic consistency, and the method comprises the steps: obtaining the three-dimensional prior from an input RGB image sequence in combination with a pre-trained multi-modal model and a three-dimensional generation model; performing semantic alignment on the three-dimensional priori and an input image based on feature similarity measurement of a visual basic model, and preliminarily estimating the pose of an object; completing initial three-dimensional reconstruction of the object according to the rough pose by using a neural radiation field; fine adjustment is carried out on the object pose based on the initial reconstruction result; and generating a high-precision three-dimensional reconstruction result through a nerve radiation field by using the optimized pose. The method has the advantages that the inherent problem of pose estimation in a hand-held object scene is solved through semantic consistency constraint; high-precision object three-dimensional reconstruction can be realized only by depending on RGB video data easy to obtain; dependence of a traditional method on complex sensor data or manual annotation is remarkably reduced, and reconstruction efficiency and practicability are effectively improved.
Owner:ZHEJIANG UNIV +1

Industrial surface defect detection method fusing large model and domain knowledge

The invention provides an industrial surface defect detection method fusing a large model and domain knowledge, and relates to the technical field of industrial detection, and the method specifically comprises the steps: designing a double-branch knowledge injection network model which comprises a backbone network, a side branch network, a texture encoder, a curvature extraction network and a detection head network, extracting general visual features by using a visual large model in the backbone network, then adding and fusing the general visual features with supplementary features extracted by the bypass network to obtain deep features, extracting texture features by using a texture encoder, and embedding texture domain knowledge into the deep features and the texture features in a cross attention mode to obtain a deep texture domain; optimizing a loss function according to the stress attention map; positioning and classifying defects by using a detection head network; and training and evaluating the double-branch knowledge injection network model. According to the technical scheme, the problems that in the prior art, the detection performance of a visual basic model in an industrial scene is reduced due to field differences, and the small sample generalization ability is weak are solved.
Owner:SHANDONG UNIV OF SCI & TECH

Tunneling rock ballast real-time identification method based on U-Net-SAM coupling driving

The invention provides a tunneling rock slag real-time identification method based on U-Net-SAM coupling driving, and the method achieves the high-precision quantization of the size and morphological parameters of rock slag through the fusion of a semantic segmentation network (U-net) and a visual basis model (Segment Anything Model, SAM). The method comprises the following steps: firstly, reasoning an input rock slag image by using a U-Net network to generate a rough rock slag identification result; then, obtaining the center-of-mass coordinate of each rock ballast from the rough result as the automatic prompt input of an SAM model, and extracting an initial mask in combination with the zero sample segmentation capability of the SAM; and carrying out post-processing optimization on the initial mask by adopting a morphological noise filtering method and an intersection-to-union ratio threshold optimization overlapped region processing strategy. According to the method, automatic segmentation and size and shape parameter extraction of the tunnel conveyor belt rock ballast image are realized, the boundary identification problem of a traditional method in a dense particle scene is effectively solved, and real-time and reliable rock machine interaction data support is provided for realizing intelligent tunneling parameter regulation and control.
Owner:CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY

Visual intelligent monitoring method and system for operation of belt conveyor

The invention relates to a visual intelligent monitoring method and system for operation of a belt conveyor. The method comprises the following steps: collecting a video stream through visual sensing equipment deployed along a belt conveyor; performing dynamic image quality evaluation and enhancement at an edge computing node, and operating a lightweight deep learning model to detect typical anomalies in real time; meanwhile, uploading the data to a cloud end, extracting high-dimensional features by using a visual basic model, and comparing the high-dimensional features with a dynamic feature distribution model to identify unknown anomalies; when an exception is found, data from the non-visual sensing module in the same time period are automatically called and analyzed to perform collaborative verification; and through a verification result, carrying out fusion decision making and generating final diagnosis and early warning information. The system comprises a visual perception module, an edge computing gateway, a non-visual sensing module and a cloud intelligent analysis platform. According to the invention, omnibearing and high-reliability intelligent monitoring and early warning of the running state of the belt conveyor are realized.
Owner:CCCC MECHANICAL & ELECTRICAL ENG

Small sample PCBA component target detection method and system based on visual basic model

The invention belongs to the field of PCBA quality detection, and relates to a small sample PCBA component target detection method and system based on a visual basic model. The method comprises the following steps: acquiring a pre-positioning prompt box from a PCBA image to be detected by using a positioning basic model, and acquiring a single-scale freezing feature of the image by using a classification basic model; expanding the single-scale features into multi-scale features for pre-positioning and classification tasks; pre-positioning prompts are screened and filtered, and query point features are obtained by using a feature alignment query sampling method; and predicting a classification result and a regression result for each query point feature according to the multi-scale positioning feature, the classification feature, the query point feature and the prompt box. According to the method, the high-precision model is quickly and finely adjusted under the condition that training data is scarce, the training cost is reduced, and the requirement for quick switching of production lines in small-batch PCBA production is effectively met.
Owner:HUAZHONG UNIV OF SCI & TECH

A method and system for zero-shot image segmentation based on object three-dimensional models

PendingCN122157261AHigh precisionEfficient automated segmentationBiological modelsKnowledge based modelsVisual BasicImage segmentation
The application relates to the field of image segmentation technology in computer vision, in particular to a zero-shot image segmentation method and system based on a three-dimensional model of an object. The method comprises the following steps: for any untrained object, a three-dimensional model of the object is rendered into a two-dimensional RGB reference image under a plurality of preset discrete viewing angles by using a graphics rendering engine, the reference image is input into a visual basic model DINOv3, a global feature vector representing semantic information of the object is extracted, and a reference feature library is constructed; a target image to be segmented is input into a visual basic model SAM2, and a plurality of candidate object masks are generated; for each candidate mask, an object image corresponding to the candidate mask is cropped from the target image, and the object image is input into the DINOv3 to extract a semantic feature vector of the candidate object; by calculating the cosine similarity between the candidate object feature vector and each template feature vector in the reference feature library, the top similar degrees with the highest values are selected, and an arithmetic mean value of the similar degrees is calculated, and the mean value is taken as the classification confidence of the candidate object mask; the class label of each candidate object mask is determined according to the confidence, and the class label of the target object and a corresponding pixel-level segmentation mask are output.
Owner:HANGZHOU HUXIYUN BAISHENG TECH CO LTD

A method for generating a sentiment data graph based on natural language understanding

ActiveCN121706786BSemantic analysisBiological modelsGraphicsVisual Basic
This invention discloses a method for generating emotional data graphics based on natural language understanding, comprising the following steps: collecting emotional text data, brand visual basic element data, and cross-cultural context description data, and processing the collected data; performing natural language understanding on the processed emotional text data; constructing a multi-domain visual semantic tensor flow that evolves over time; constructing a cross-cultural intent field and binding the parameters of the cross-cultural intent field to the multi-domain visual semantic tensor flow; generating a preliminary emotional graphic structure sequence through a generative visual network; dynamically adjusting the preliminary emotional graphic structure sequence; and outputting it to different media scenarios to obtain emotional data graphics. This invention integrates natural language understanding and multimodal generation technologies to achieve emotion-semantic driven graphics generation, possessing the advantages of semantic accuracy, cultural adaptation, and visual coherence.
Owner:DALIAN POLYTECHNIC UNIVERSITY

Cross-domain small sample wideband signal detection and identification method based on visual foundation large model

PendingCN122637056AVisual BasicAlgorithm
The application provides a cross-domain small sample wideband signal detection and recognition method based on a visual basic model, and relates to the cross technical field of electromagnetic signal processing and computer vision. The application reduces the interference of background noise in the wideband time-frequency graph on the candidate area by performing double condition screening based on target confidence and aspect ratio features on the preliminary candidate frame, and combining the field of view expansion in the frequency axis direction, and improves the incomplete coverage problem of the slender signal candidate frame. By extracting multiple intermediate layer features and the final layer feature map of the visual basic model for fusion in the channel dimension, the problem that the local texture information in the wideband time-frequency graph is smoothed or covered in the layer-by-layer abstraction process is solved. The parameters of the visual segmentation model and the visual basic model are kept frozen, and only the small sample fine tuning of the front region proposal network is performed, so that the calculation overhead caused by full fine tuning of the visual basic model during cross-domain small sample adaptation is avoided, and the method is suitable for deployment on devices with limited computing resources.
Owner:NORTHEASTERN UNIV CHINA

Cross-domain small sample semantic segmentation method and device, equipment and medium

The invention discloses a cross-domain small sample semantic segmentation method and device, equipment and a medium. According to the cross-domain small sample semantic segmentation method, a hierarchical data set is constructed through a controllable style offset synthesis image, and a style feature extraction and weight generation network is trained; a dynamic weight is generated based on the difference between the target domain image style representation and the source domain reference, feature layer level modulation is performed on the visual basis model encoder features, and the feature alignment problem is accurately solved; image semantic information is fused and supported through a memory attention mechanism, so that the problem of small sample semantic sparsity is effectively relieved; meanwhile, the synthetic data accurately controls style offset through a stable diffusion model, and uncontrollability of traditional generative data expansion is avoided. According to the method, the cross-domain small sample semantic segmentation performance is remarkably improved, and the problems of prompt mismatching, feature dislocation and out-of-control data expansion caused by domain offset are solved.
Owner:SHENZHEN UNIV

Fine adjustment method, system and equipment for visual basic model and storage medium

The invention discloses a fine tuning method, a fine tuning system and fine tuning equipment for a visual basic model, and a storage medium, which are corresponding schemes, and the related scheme aims to solve the problems of overlarge GPU video memory consumption, high quantitative perception fine tuning overhead and the like in the fine tuning and deployment process under the condition of limited resources in the existing parameter efficient fine tuning scheme. According to the scheme, through a sub-network adapter quantization perception fine tuning and block-level activation quantizer fine tuning strategy, video memory consumption in a fine tuning stage is remarkably reduced, and meanwhile, the calculation complexity consistent with that of an original backbone network in a reasoning stage is kept; therefore, a feasible and efficient solution is provided for low-resource deployment of the visual basic model in diversified downstream tasks.
Owner:UNIV OF SCI & TECH OF CHINA

Information visualization method and system applied to intelligent building carbon emission monitoring

The invention provides an information visualization method and system applied to intelligent building carbon emission monitoring, and the method comprises the steps: firstly obtaining an initial carbon emission data set of different functional regions of an intelligent building in a continuous monitoring period, and then carrying out the feature extraction of the initial carbon emission data set; generating a time sequence fluctuation feature set and a spatial distribution correlation feature set, then inputting the time sequence fluctuation feature set and the spatial distribution correlation feature set into a machine learning network, generating a visual basic data set containing a carbon emission evolution trend and a regional difference feature, and according to a visual coding rule, determining the carbon emission evolution trend according to the visual basic data set. According to the method, the characteristic components are mapped into graphical representation parameters, and finally, visual graph data containing a time axis carbon emission fluctuation thermodynamic diagram and a building space carbon emission distribution color block diagram are generated based on the graphical representation parameters, and are presented in real time through an interactive display interface, so that management personnel can quickly and accurately master the carbon emission condition of a building, and the construction efficiency is improved. And abnormal carbon emission regions and evolution trends can be found in time.
Owner:CHINA CONSTR CARBON TECH CO LTD +1

An agent method and system for continuous evolution of robotic job skills

The application discloses an agent method and system for continuous evolution of robot operation skills, comprising a skill sharing semantic rendering module, a skill sharing representation distillation module and a skill specific evolution planner. In view of long-term challenges such as 3D scene representation and human task learning faced by robots when adapting to new sequence tasks in actual scenes, the application can continuously learn new 3D scene semantics and robot operation skills from skill sharing and skill specific attributes. The skill sharing semantic rendering module and the skill sharing representation distillation module effectively learn 3D scene semantics by means of neural radiance fields and visual basic models, solve the problem of 3D scene representation neglect, and use the skill specific evolution planner to decouple skill knowledge through latent and low-rank space, and continuously embed new skill specific knowledge. The application also designs a robot operation benchmark, and experiments show that the method significantly improves the performance in robot operation tasks.
Owner:SOUTH CHINA UNIV OF TECH

Complex structure wheel forging rolling whole process digital design and performance prediction method

The invention discloses a digital design and performance prediction method for the whole forging and rolling process of a wheel with a complex structure, and belongs to the technical field of metal plastic forming and digital simulation. In order to overcome the defects of an existing method, an integrated forming module is designed; the invention provides a parametric modeling method, based on Visual Basic and SolidWorks secondary development, wheel parameters in an Access database are accessed, and three-dimensional modeling and assembling of all molds and blanks are automatically completed; a structure performance prediction method is provided, a parameterized model is automatically imported into finite element software, a microstructure model and a cellular automaton are integrated through secondary development, macro-micro coupling calculation is carried out, and microstructure distribution after the wheel is formed is predicted. According to the method, full-process digitization and automation from wheel design parameter input to final structure performance prediction are achieved, the research and development efficiency and the forming quality prediction precision of the complex-structure wheel are remarkably improved, and a core technical support is provided for optimal manufacturing of rail transit key components.
Owner:TAIYUAN UNIVERSITY OF SCIENCE AND TECHNOLOGY

Visual question-answering method and device based on multiple modes, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to the fields of financial science and technology and medical health, and discloses a multi-modality-based visual question and answer method, device, equipment and medium, and the method comprises the steps: obtaining an input image and a question, and inputting the image and the question into a visual language coding model for coding to generate a multi-modality feature; inputting the multi-modal features into an answer reasoning module to generate answers and deep reasoning features corresponding to the answers; inputting the multi-modal features into a basic principle generation module to generate guide features, and inputting the guide features into a large language model to generate a text basic principle; extracting features of the text basic principle to obtain context features, and generating a visual basic principle through an object detector according to the context features, the image and the depth reasoning features; and taking the visual basic principle, the text basic principle and the answers as final interpretable answers. And the transparency and the credibility of the visual question-answering system are effectively improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Method and system for weed seed identification based on DEIMv2 and identification special model architecture

ActiveCN122392034BVisual BasicFeature extraction
The application discloses a weed seed identification method and system based on DEIMv2 and a special model architecture for identification, wherein the model architecture comprises: a multi-scale space adapter which extracts spatial features and performs multi-scale down-sampling on a target image to output three spatial feature tensors with descending resolutions; a visual basic model pre-trained in a self-supervised manner which encodes deep semantic features of the target image and outputs multi-scale deep semantic feature tensors; a hybrid encoder which aggregates the spatial feature tensors and the deep semantic feature tensors at the same scale to obtain first, second and third aggregation feature tensors with descending scales; a feature enhancement path is: the first aggregation feature tensor is fused with the second aggregation feature tensor after spatial channel down-sampling, and then fused with the third aggregation feature tensor after spatial channel down-sampling to output an enhanced feature tensor; and a DEIM decoder and a detection head, wherein the decoder receives the enhanced feature tensor and decodes.
Owner:COMPREHENSIVE TECH CENT FOR INSPECTION & QUARANTINE OF ZHANGJIAGANG ENTRY EXIT INSPECTION & QUARANTINE BUREAU +1

A multi-modal knowledge guided transmission line defect detection method

The embodiment of the application discloses a kind of multi-modal knowledge guided transmission line defect detection methods, including steps: constructing transmission line image-text data set, through image-text contrast learning pretraining, the relevant multi-modal knowledge of transmission line contained therein is introduced into visual basic model CLIP;The embedding matrix of CLIP model pre-encoded carrying the relevant multi-modal knowledge of transmission line is introduced into the decoder of Deformable DETR detector, to realize the detector training of multi-modal knowledge guidance;Residual structure is introduced into the decoder of Deformable DETR, to alleviate the problem that the multi-modal knowledge introduced is gradually forgotten in the attention calculation of decoder multi-layer;Design a kind of pseudo-label distillation objective function, with multi-modal knowledge constraint optimization direction in the detector training process, realize more stable training effect.The application uses multi-modal knowledge to guide the training process of transmission line defect detector, effectively improves the precision of defect detection, with good robustness and stability.
Owner:NORTH CHINA ELECTRIC POWER UNIV

A zero-shot object instance segmentation system and method under a robot environment

The present invention discloses a zero-sample object instance segmentation system and method in a robot environment, which belongs to the field of robot visual perception technology. First, the depth image is preprocessed with Viridis color mapping, and then the SAM (Segment Anything Model) model is used to generate initial object mask candidates; at the same time, the pre-trained ViT (Vision Transformer) is used as the feature description model of the scene to process the image, extract the attention map of the last layer, and construct a feature weighting mechanism based on the information entropy of the attention map; then the similarity matrix of the image block and the background block is calculated to remove the non-object mask, and the K-Medoids clustering algorithm is used to obtain representative sampling points for each remaining object mask candidate. Finally, these points are input into the SAM model as prompt information to obtain accurate object instance segmentation results. The present invention makes full use of the zero-sample generalization ability of the visual basic model and can achieve accurate segmentation of unseen objects without additional training, with good versatility and practical value.
Owner:YANSHAN UNIV

Image tampering detection method based on visual basic model

The invention provides an image tampering detection method based on a visual basic model, which comprises the steps of cross-view automatic prompt generation, feature extraction from various complementary views and fusion into prompts, and manual prompt dependence elimination. Selecting an optimal prompt based on the minimum segmentation loss through an optimal prompt selection module; utilizing a cross-view consistency module to enhance view prompt consistency with consistency loss; multi-view features are aligned and fused through a cross-view feature perception module; fusing the features and the optimal prompt through a prompt mixing module to generate final prompt embedding; and finally, inputting the fusion prompt and the feature into an SAM mask decoder to realize tampering detection and positioning. According to the method, the visual basic model is applied to tampering detection, and the performance is improved through multi-view processing. Experimental results show that the method significantly improves model generalization, robustness and tampering positioning precision.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

A semantic matching method based on a large-scale pre-training model visual feature

The application discloses a semantic matching method based on a large-scale pre-training model visual feature, comprising the following steps: acquiring an image data set, which comprises a target image and a reference image corresponding to a semantic; constructing a corresponding semantic matching network based on a visual basic model framework, wherein the semantic matching network comprises a feature extraction module, an interlaced perception module, a matching module and an up-sampling module; training the semantic matching network by using the image data set to obtain a semantic matching model used for image semantic matching; and inputting the reference image, a point to be matched on the reference image and the target image into the semantic matching model to output a matched point on the target image. The method provided by the application can effectively improve the accuracy of a semantic matching task and provide better services for downstream visual tasks.
Owner:ZHEJIANG UNIV

RGB-D video salient target detection method based on depth-guided adaptive query

The invention relates to an RGB-D video salient target detection method based on depth guidance adaptive query, which comprises the following steps: firstly, constructing a parallel adapter structure based on depth guidance, and adopting a parallel jump structure to remarkably reduce the video memory overhead brought by gradient return during fine tuning, so that an SAM can automatically segment a salient target without any prompt; and meanwhile, multi-modal fusion of RGB-D features is realized by taking the depth map as guidance. And secondly, providing a query-driven time sequence memory module, replacing a huge Memory Bank in the SAM2 with frame-level query and video-level query, and realizing lightweight time sequence modeling. By introducing and improving a visual basic model structure, efficient and accurate video salient target segmentation is realized under the condition of no artificial prompt, and meanwhile, the video memory consumption and the calculation complexity of the model are remarkably reduced.
Owner:HANGZHOU DIANZI UNIV

Virtual object special effect logic editing method and device, equipment and medium

This invention discloses a method, apparatus, device, and medium for editing virtual object special effects logic. The method includes: receiving a configuration file and parsing it to obtain a configuration database; determining a selected object corresponding to selected information in an imported file; performing logical configuration on the selected object according to the configuration database to generate visual basic logical information; adjusting the parameters of the basic logical information according to parameter modification information; obtaining candidate special effects actions for the selected object from the configuration database; configuring special effects actions on the adjusted basic logical information according to action selection information; editing the configuration information in the configuration database; and generating a configuration editing file. This method can parse the configuration file to obtain a configuration database and perform logical configuration on the selected object, generating visual basic logical information for efficient parameter adjustment and special effects action configuration, thus improving the efficiency and accuracy of editing the characteristic logical information in the configuration file.
Owner:HANGZHOU WIZARD GAME TECH CO LTD

Machine learning model input monitor

The invention relates to a method, in particular a computer-implemented method, for monitoring the performance of a trained machine learning model, a corresponding monitoring system, a computer program and a computer-readable medium. The method comprises the following steps: providing a trained visual basic model; providing a training data set for training the machine learning model; for each training image in the training data set, using the visual basic model to determine a training data feature vector, and determining the distribution of the training data feature vector; receiving an image depicting a scene; determining an image feature vector of the image by using the visual basic model; calculating a log-likelihood value of the distribution of the image feature vector to the training data feature vector; if the image differs from the distribution of the training data feature vectors based on the analysis of the log-likelihood values, an alert is provided.
Owner:CONTINENTAL AUTONOMOUS MOBILITY GERMANY GMBH

A spatial target recognition method based on bimodal fusion

The application discloses a space target recognition method based on bimodal fusion, and belongs to the technical field of space target recognition. The method comprises the following steps: processing a space target image to generate a structured text description, and inputting the space target image and the structured text description into a visual basic model to obtain two-dimensional classification logic; performing farthest point sampling and dynamic updating on three-dimensional point cloud matched with the space target image to obtain an updated center point set of the three-dimensional point cloud; performing k-NN search, local aggregation and global pooling processing on the updated center point set to obtain global features with semantic coherence; and using the two-dimensional classification logic and the global features to obtain a final space target recognition result of the space target image to be recognized. Through cooperative processing of two-dimensional images and three-dimensional point cloud data, and in combination with field adaptation and non-parametric feature matching of a visual basic model, the application can realize high-precision and high-robustness space target recognition.
Owner:XIDIAN UNIV

A frequency-space dual-domain joint fine-tuning visual base model network method for remote sensing image domain generalization semantic segmentation

The application discloses a kind of frequency-space dual-domain joint fine-tuning visual basic model network methods for remote sensing image domain generalization semantic segmentation, which enhances the consistency of cross-domain similar features by adaptive frequency selection in the middle features of the frequency domain fine-tuned model, and improves the discriminant ability of the feature cluster boundary in the deep features of the spatial domain fine-tuned model to achieve robust domain generalization semantic segmentation performance. Through four experimental settings of three public datasets Potsdam, Vaihingen and LoveDA, the mIoU index of the present method is improved by 1.64% and 1.09% on average compared with the most advanced method, and the model parameter amount is only 4.22M, which balances the accuracy and computational efficiency.
Owner:BEIJING INST OF TECH