Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

36 results about "Visual Word" patented technology

Visual words, as used in image retrieval systems, refer to small parts of an image which carry some kind of information related to the features (such as the color, shape or texture), or changes occurring in the pixels such as the filtering, low-level feature descriptors (SIFT, SURF, ...etc.).

Method and device for positioning and mapping industrial automobile crane based on multi-sensing fusion

The invention discloses an industrial automobile crane positioning and mapping method and device based on multi-sensing fusion, and relates to the field of machine sensing, and the method comprises the following steps: initializing a pose and offset of an inertial measurement unit by using an input laser point cloud and a pre-integration result of the inertial measurement unit; correcting laser point cloud motion distortion by using an inertial measurement unit, extracting a dynamic object point cloud, and synchronously extracting laser features and visual features with depth; calculating a key frame pose in real time by minimizing a visual re-projection error and an inertial measurement unit pre-integration error, and performing frame-to-map matching by taking the key frame pose as an initial value to further optimize the pose; a visual word bag model is utilized to quickly retrieve a closed-loop candidate frame, after laser passes geometric matching to verify a closed loop, a laser radar factor, a visual factor, an inertial measurement unit factor and a closed-loop factor are fused, a global factor graph is constructed and optimized, and a pose and point cloud map is updated. According to the method, the environment sensing capability of the industrial automobile crane can be improved, and the sensing accuracy and robustness are improved.
Owner:NANKAI UNIV +1

Wall-climbing robot positioning and path planning method based on unmanned aerial vehicle cooperation

The invention discloses a wall-climbing robot positioning and path planning method based on unmanned aerial vehicle cooperation, and the method comprises the steps: constructing and updating a three-dimensional map based on the point cloud and image data of a laser radar and a camera; based on the three-dimensional map, the two-dimensional code of the wall-climbing robot is identified by using the unmanned aerial vehicle, and the accurate positioning information of the wall-climbing robot is obtained by combining self-positioning and cross validation and fusion correction of the ground control unit; updating a map by synchronizing sensing data, and optimizing a climbing track based on the obtained positioning information and a global planning path of the ground control unit; for positioning information and environment information of the wall-climbing robot, the unmanned aerial vehicle adopts a trajectory prediction control algorithm to adjust the attitude, and meanwhile, sudden obstacles are avoided in real time through path planning. According to the invention, through laser vision fusion SLAM and unmanned aerial vehicle collaborative operation, the environment perception capability is enhanced, and through combination of a dynamic map and visual word bag optimization, the functions of operation map construction, accurate positioning, path planning, autonomous navigation obstacle avoidance and the like of the wall-climbing robot are realized.
Owner:HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES

Tree species identification method based on visual word bag representation and state space modeling

The invention discloses a tree species identification method based on visual word bag representation and state space modeling, and relates to the technical field of tree species cross section microscopic image identification, and the method is characterized in that the method comprises the following steps: S1, data set construction; s2, extracting local features by using a visual word bag model; s3, extracting global features by using the state space model; s4, performing multi-level feature fusion; s5, data oversampling and classifier training are carried out; and S6, identifying tree species. The technical problem to be solved by the invention is to provide a tree species identification method based on visual word bag representation and state space modeling, a state space model is introduced to model a feature sequence of a tree species image, a long-term dependency relationship of the feature sequence is mined, and global feature information is extracted. Through fusion of local and global features, multi-level feature representation with strong discrimination capability is constructed. Samples are balanced by adopting a synthetic minority class oversampling technology so as to improve the discrimination capability of a support vector machine classifier.
Owner:SHANDONG JIANZHU UNIV

Visual text generation method, model training method, intelligent agent and device

The invention provides a visual text generation method, a model training method, an intelligent agent and a device, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like, and can be applied to scenes such as content generation based on artificial intelligence. According to the specific implementation scheme, language structure information of visual characters in an image is trained into a discrete submerged space of the image, an autoregression generator with a structure understanding capability is constructed, image discretization representation can be generated, decoding processing is carried out on the image discretization representation, a reconstructed image comprising the visual characters is obtained, and the visual characters are obtained. And determining a loss function based on the image tag and the reconstructed image in the image training data and the language structure information of the visual text in the reconstructed image, and carrying out model training based on the loss function to obtain a trained visual text generation model.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Method for accelerating multi-modal large language model reasoning based on speculative decoding

PendingCN120525058ABiological modelsInference methodsBasic languageData set
The invention discloses a method for accelerating reasoning of a multi-modal large language model based on speculative decoding, the input of the multi-modal large language model is composed of a system lexical element, a visual lexical element, an instruction lexical element and an output lexical element, and the method comprises the following steps: S1, extracting a predetermined number of samples from a ShareGPT data set which is finely adjusted by using an LLaMA model instruction, pre-training an Eagle small model, and obtaining a pre-trained Eagle small model; enabling the Eagle small model to establish a basic language generation capability in a plain text environment; s2, extracting a predetermined number of samples in a ShareGPT data set finely adjusted by an LLaVA model instruction, and migrating the capability of an Eagle small model by using a transition mode; s3, after the draft model is subjected to two-stage training, on the basis of the improved draft mechanism design, differential processing is carried out according to the characteristic difference between the text and the visual mode; according to the method, the performance of speculative decoding in a multi-modal task can be improved, the acceptance length of each verification of the target model is improved, and respective adaptation of the text lexical elements and the visual lexical elements in the draft model is realized.
Owner:XIAMEN UNIV

Training method of image encoder, image processing method, electronic device, storage medium and program product

The embodiment of the application provides a kind of training method of image encoder, image processing method, electronic equipment, storage medium and program product, it is related to image processing technical field, the training method of image encoder includes: training image is input image encoder, visual word symbol for representing image feature is extracted by the image encoder;Clothing text word symbol is generated for training image;The clothing text word symbol is used to describe the clothing feature of the person in the training image;The model parameter of the image encoder is iteratively trained using a combination loss function;The combination loss function is used to make the visual word symbol of the same person of multiple training images be smaller, and make the visual word symbol and the clothing text word symbol of the same training image be larger difference.It reduces the degree of dependence on visual information by introducing text modal, increases the precision of image encoder to identify the same person under different clothing, improves the precision and robustness of person re-identification.
Owner:PEKING UNIV

Detecting fields in document images

A method of detecting fields in document images includes: receiving, by a processing device, a codebook comprising a set of visual words, each visual word corresponding to a center of a cluster of local descriptors, wherein each local descriptor is associated with a respective keypoint region of a first set of document images; calculating, based on a second set of document images, for each visual word of the codebook, a respective frequency distribution of a field position of a specified field with respect to the visual word; loading a document image for extraction of target fields; and detecting fields in the document image based on the calculated frequency distributions.
Owner:ABBYY DEVELOPMENT INC

Electric power vision large model multi-scale semi-supervised target detection method and system

The invention discloses an electric power vision large model multi-scale semi-supervised target detection method and system, and the method comprises the steps: training a multi-scale vision word segmentation device based on a vector quantization knowledge distillation framework, and constructing an electric power semantic enhanced multi-scale vision codebook, the multi-scale visual codebook comprises a top-layer codebook used for encoding a global structure of the power equipment and a bottom-layer codebook used for encoding defect local features of the equipment; based on the multi-scale visual codebook, adopting a mask image modeling task to pre-train a visual large model on an unmarked power image; performing semi-supervised fine tuning on the visual large model by using the labeled electric power image data and the unlabeled electric power image data; using an SMLS fine tuning strategy to update the visual large model parameters; and inputting a to-be-detected power image into the updated visual large model, and outputting a detection result of the multi-scale target. According to the method, the problems of scarcity of annotation data, model semantic gaps and high calculation complexity in an electric power scene in the prior art are solved.
Owner:STATE GRID ELECTRIC POWER RES INST +1

Training methods for scene classification models, remote sensing image scene classification methods and equipment

This invention provides a training method for a scene classification model, a remote sensing image scene classification method, and an apparatus. The specific implementation includes: acquiring multiple first image blocks and their respective sample labels; using a feature extraction network to extract features from each first image block to obtain local features; generating a visual word histogram for each first image block based on its local features; using a first classification network and a second classification network based on the visual word histogram to obtain a first classification result and a second classification result corresponding to each scene to be classified; determining a first loss value for the first classification result and a second loss value for the second classification result based on the correlation between the first classification result, the second classification result, the sample labels, and the multiple scenes to be classified; and adjusting the parameters of the first classification network and the second classification network according to the first loss value and the second loss value, respectively.
Owner:BEIJING DATA INTELLIGENCE INFORMATION TECH CO LTD

Method for hallucination detection and suppression of multi-modal large model based on attention time difference

The application discloses a hallucination detection and suppression method based on attention timing difference of a multimodal large model, which comprises the following steps: inputting text word elements of a prompt word and visual word elements of an image into a multimodal large model to obtain an internal attention graph in a decoding stage of the multimodal large model; then, calculating an attention proportion of the visual word elements in a current generation moment and performing difference with the proportion in a previous moment; if the difference exceeds a set threshold, it is determined that the visual word elements are related; when the visual related word elements are identified, a second forward propagation with visual enhancement is performed to obtain more accurate output; if not, the next generation is directly entered. According to the application, the visual related word elements in text generation are identified and refined through attention timing difference, twice forward propagation is adopted, the visual attention of the second forward propagation is enhanced based on the visual attention graph of the first forward propagation, and hallucination can be identified and corrected without reducing the language expression ability.
Owner:HANGZHOU DIANZI UNIV

A semantic selection and quota control method and system based on user vocabulary state

PendingCN122311131AData miningTarget text
This invention discloses a semantic selection and quota control method and system based on user vocabulary state. The method includes: acquiring target text and identifying word-level candidate sequences in the text, where the candidate sequences preserve the structural positional relationships of words within the text; constructing a user vocabulary state space, modeling the vocabulary selection process as a path construction process under the constraints of the state space Ω, where candidate words serve as path nodes, and whether a node enters the path depends on its state label and constraint function; dividing the candidate sequences based on the state space Ω to obtain a set of known words and a set of unknown word candidates; during the path construction process, imposing constraint functions on the unknown word candidate set, and selecting words from it to generate a new word set; merging the new word set with the known word set to generate a visual word path, and generating semantically enhanced text data for computer display or output based on the visual word path.
Owner:CHUANGZHI YUNWEI (BEIJING) TECHNOLOGY CO LTD

A method and related equipment for semantic understanding of long-term traffic videos

This application provides a method and related equipment for semantic understanding of long-term traffic videos, relating to the fields of intelligent transportation and computer vision technology. This application reconstructs a target long-term traffic video sequence into a target sparse traffic visual word sequence involving dynamic traffic visual concepts. Then, it performs cross-timescale sparse attention feature learning on the target sparse traffic visual word sequence across multiple traffic state perception dimensions. Based on the assigned desired video semantic understanding task, it fuses the local attention output feature sequences of each of the learned traffic state perception dimensions to obtain the target traffic video semantic feature sequence, minimizing redundant information interference during the traffic video semantic understanding process. Finally, it invokes a large language model to perform cross-modal attention interaction reasoning based on the target traffic video semantic feature sequence and the desired video semantic understanding task, effectively balancing the efficiency and accuracy of video semantic understanding.
Owner:BEIHANG UNIV

A core subgraph-based training-free multi-modal large language model fine-grained positioning method

PendingCN122451086ALinguistic modelAlgorithm
The application belongs to the technical field of computers, and particularly relates to a core subgraph-based training-free multi-modal large language model fine-grained positioning method. The method comprises the following steps: 1) using a pre-trained multi-modal large language model to respectively perform feature coding on input images and texts, and respectively dividing the images and texts into a plurality of word units to construct visual and text word unit vector sequence sets; 2) constructing a visual semantic graph structure according to the semantic similarity relationship between visual word units, taking each visual word unit as a graph node, and taking the similarity between the nodes as an edge weight, so as to form a normalized adjacency matrix; 3) calculating the importance score of each visual word unit, and taking a region specified by a user in an arbitrary shape as a core subgraph; 4) performing a visual attention diffusion process in the visual semantic relationship graph based on the core subgraph, and determining a subgraph according to the importance distribution after diffusion; and 5) inputting the screened subgraph and the coded text word unit into a large language model to generate an answer through decoding.
Owner:NANKAI UNIV

Auto-regression image generation method and device based on multiple solution terminals

The invention provides an autoregression image generation method and device based on multiple solution wharfs. The method comprises the following steps: adding a plurality of parallel solution wharfs on an initial autoregression model; inputting a historical image set serving as a training sample set into a plurality of parallel solution wharfs for training to obtain a plurality of trained parallel solution wharfs; inputting a to-be-recognized image into the plurality of trained parallel solution wharfs, and predicting the positions of the plurality of visual lexical elements in the model reasoning process by using the plurality of trained parallel solution wharfs to obtain a target prediction result; and based on the target prediction result, using a speculation decoding strategy to perform decoding operation on the plurality of visual lexical elements to obtain a target autoregression image corresponding to the plurality of visual lexical elements. Under the condition that the autoregression image generation quality and the diversity pre-training model are ensured, the autoregression text is accelerated to generate the image, the image generation quality is improved, the training resource consumption of the autoregression model is reduced, and the calculation cost is reduced.
Owner:BEIJING AUTONAVI YUNMAP TECH CO LTD

SLAM loopback detection method and system based on pipeline topology constraint

The invention discloses a pipeline topology constraint-based SLAM loopback detection method and system, and belongs to the technical field of robot synchronous positioning and mapping. Comprising the following steps: acquiring a prior topological map of a pipeline scene; retrieving loopback candidate frames through a visual word bag model; the method is characterized in that SLAM estimation poses of a current frame and a candidate frame are mapped to nodes of a topological map, and a topological path distance between the two nodes is calculated; and verifying and screening the loopback candidates based on a consistency comparison result of the distance and a preset threshold value. The device is mainly used for providing high-reliability loopback detection capacity for the oil and gas pipeline inspection robot. According to the method, the prior topology constraint is introduced as a strong filter, false loops can be effectively eliminated from the source, and the robustness and the positioning precision of the SLAM system in a structured industrial environment are remarkably improved.
Owner:SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI +1

Synchronous positioning and mapping method and system based on deep learning

The invention discloses a deep learning-based synchronous positioning and mapping method, which comprises the following steps of: extracting image features by using a deep learning model to obtain a feature point probability tensor and a descriptor tensor; adaptively screening feature points in the feature point probability tensor according to the inter-frame feature relationship and the intra-frame feature relationship, and generating a screened feature point set; constructing a graph neural network model by using the screened feature point set and the descriptor tensor; obtaining a feature point matching relation according to an output result of the graph neural network; performing binarization processing on the depth feature descriptors of the floating-point number, and training a visual word bag by using the feature descriptors after binarization processing; performing closed-loop detection by using the trained visual word bag, and retrieving similar key frames; and carrying out pose optimization and map construction by utilizing the retrieved similar key frames and feature matching results. According to the invention, a more accurate scene identification effect is realized, and real-time, robust and accurate pose estimation and map construction of the hybrid SLAM system are also realized.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

Fine-grained image retrieval method based on compact representation modeling and semantic label guidance

The present application relates to a fine-grained image retrieval method based on compact representation modeling and semantic label guidance, and belongs to the technical field of fine-grained image analysis, solves the problem that it is difficult to efficiently mine significant and compact intermediate features of images in the prior art, comprising: step S1: preparing a pre-processed image data set; step S2: constructing a feature extraction network to extract visual word vectors; step S3: generating a hash coding center; step S4: constructing a compact coding mapping unit; step S5: training the constructed feature extraction network and compact coding mapping unit to obtain a trained feature extraction network and compact coding mapping unit; step S6: fine-grained image retrieval, matching the obtained to-be-detected code with the codes in the retrieval code library, and the sample with the smallest distance is the final retrieval result.
Owner:BEIHANG UNIV

Information retrieval method, device, storage medium, and program product

PendingCN122332597AData setData mining
This invention provides an information retrieval method, device, storage medium, and program product. The method includes: acquiring a representation model trained on a training dataset, wherein different retrieval tasks correspond to different data modality combinations in the training data pairs, and the text data used to describe image data in the image-text pairs includes visual words that meet a set density condition; acquiring retrieval input data and multiple retrieved data corresponding to the target retrieval task; determining the embedding representation vectors corresponding to the retrieval input data and the multiple retrieved data respectively through the representation model; and selecting the target retrieved data from the multiple retrieved data based on the similarity between the two. The training dataset uses dense multimodal knowledge and training data from various retrieval tasks, improving the multimodal alignment capability of the representation model, enabling it to adapt to different types of retrieval tasks, and improving the accuracy of retrieval results.
Owner:ALIBABA CLOUD COMPUTING CO LTD

High-resolution remote sensing image cloud computing optimization distributed retrieval method

According to the high-resolution remote sensing image cloud computing optimization distributed retrieval method, on the basis of a self-defined MapReduce architecture IO method, a Hadoop-based high-resolution remote sensing image distributed retrieval architecture with offline library building and online retrieval is realized; a target detection method is integrated into a retrieval architecture, a method and parameters of an optimal alternative region extraction effect are determined through comparison and analysis of multiple homogeneous region extraction methods and different parameter settings and calculation of time consumption and result precision, and target positioning during retrieval is realized; the method comprises the following steps of: customizing a remote sensing image reading format and an image partitioning mode by adopting a GDAL, extracting various features of an original remote sensing image in a Map stage, searching an optimal feature combination, generating a visual word bag model through a clustering algorithm based on MapReduce, constructing a high-resolution remote sensing image database and a feature database based on Hadoop, carrying out distributed storage on image data and feature data, and storing the image data and the feature data. Retrieval acceleration is realized through parallel computing; the high-resolution remote sensing image retrieval speed is high, and the precision is high.
Owner:王飞

A new method for open-vocabulary semantic segmentation of remote sensing images

The application discloses a new method for remote sensing image open vocabulary semantic segmentation. The method comprises the following steps: a dense feature calibration stage, which improves the spatial reliability of low-resolution visual words by abnormal feature repair and semantic correlation attention mechanism; an arbitrary resolution feature up-sampling stage, which reconstructs low-resolution features into high-resolution dense representations by image-guided localized cross-attention guided by input image structure clues; and a coordination aggregation stage, which aggregates the features extracted from multiple geometrically transformed views after inverse transformation alignment to suppress perspective-related fluctuations. The application does not require remote sensing pixel labeling, can be used immediately after migration among the CLIP series backbone, and is suitable for open vocabulary pixel-level recognition in multiple types of remote sensing scenes such as satellite images and unmanned aerial vehicle aerial images.
Owner:BEIJING UNIV OF POSTS & TELECOMM

A key frame-based mobile robot efficient closed loop detection method

The application discloses a kind of mobile robot high-efficiency closed loop detection methods based on key frame, comprising: obtaining image frame;With visual odometer, the key point and descriptor of current image frame are extracted, and the pose estimation of current image frame and its adjacent image frame is carried out;With semantic segmentation network, current image frame is converted into semantic graph, and semantic segmentation result is obtained;The first image frame collected when mobile robot runs is recorded as key frame, and the rest key frame is extracted according to pose estimation and semantic segmentation result;The key point and descriptor of each key frame are input into BoW model and converted into corresponding visual word vector, and visual word vector is stored in historical dataset;Loop closure detection is carried out on current image frame to obtain loop closure detection result;According to loop closure detection result, switch loop closure detection state.The method can adaptively switch detection state, reduce computational load, and effectively improve the key frame selection result and loop closure detection efficiency of system on the basis of ensuring loop closure detection accuracy.
Owner:ZHEJIANG UNIV OF TECH

A visual word pruning method and device based on cross-layer attention evolution

The application discloses a visual word pruning method and device based on cross-layer attention evolution, and solves the technical problem that the existing dynamic pruning method based on a large model seriously limits the scalability of the dynamic pruning framework in a concurrent scene. The method comprises the following steps: acquiring an input image to be inferred and user query text, completing image block coding and cross-layer attention weight collection through a pre-trained visual encoder, outputting a complete spatial visual word sequence and a shallow and deep anchor layer full attention head attention weight matrix; obtaining an anchor word index set through double-anchor keyword index screening, and completing diversity complementary selected elements and sequence regularization, and outputting a simplified visual word sequence; finally, combining a pre-trained text segmenter, a multi-modal projection layer and a large language model backbone network to perform multi-modal joint inference, and outputting a multi-modal inference response text.
Owner:SUN YAT SEN UNIV

Multi-view target detection method, device, computer equipment and storage medium

An embodiment of the present invention proposes a multi-view target detection method, apparatus, computer equipment and storage medium, which belong to the field of computer vision. The method includes: obtaining multiple local image blocks of an image to be tested, and image block features of each local image block, and using a preset random forest model based on each image block feature to divide the multiple local image blocks into multiple visual words, thereby calculating the voting score of each visual word for each candidate center position of the image to be tested, obtaining the voting combination weight of each candidate center position with respect to each visual word, and then calculating the total score of each candidate center position based on the voting combination weight and the voting score, and determining the target center position based on the total score of each candidate center position. The method can fully consider the perspective factor, thereby reducing the influence of interference such as target occlusion, deformation and angle change, and improving the accuracy of target detection.
Owner:NANJING DANIU INFORMATION TECH CO LTD

Industrial bearing data enhancement method based on visual words and diffusion model

The invention belongs to the technical field of bearing fault detection, and particularly relates to an industrial bearing data enhancement method based on visual words and a diffusion model. The method realizes data enhancement based on a more stable conditional diffusion probability model (DDPM) and visual words, captures features of an original continuous wavelet transform image through the visual words, converts the features into feature vectors, expands the feature vectors through the DDPM, takes the feature vectors as input through the DDPM, and outputs a synthesized wavelet transform image. According to the method, the generation model independent of the discrimination model is designed, the problem that the Gan model is difficult to train is solved by using the diffusion model, and the mapping relation in the generation process is enriched through the feature vectors generated by the visual words, so that the types of the generated data are richer.
Owner:BEIJING JIAOTONG UNIV +1

Synchronous positioning and mapping method and system for multi-source information fusion

The invention provides a multi-source information fusion SLAM (Simultaneous Localization and Mapping) algorithm and system. Vision, laser radar and IMU (Inertial Measurement Unit) data are deeply fused. A system framework comprises data preprocessing, pose estimation, back-end optimization and a hardware platform. The data module preprocesses the laser point cloud; an improved ICP laser odometer and a visual odometer (combining optical flow tracking and angular point detection) are adopted for pose estimation, and error state Kalman filtering fusion is carried out. And back-end optimization is based on laser point cloud and visual word bag model design loopback detection, accumulative errors are reduced, and a color point cloud map is generated. A hardware platform is carried on an unmanned aerial vehicle and an unmanned vehicle, a multi-sensor scanner is constructed, sensor calibration is carried out, and remote networking and ROS distributed configuration are adopted. The method is high in positioning precision, high in robustness and centimeter-level in mapping precision in a complex environment, information of multiple sensors is effectively fused, the generated color RGB point cloud map is close to a real space, construction of a digital twinborn model is facilitated, and AR / VR technology development is promoted.
Owner:项朝阳

Detecting fields in document images

A method of detecting fields in document images includes: receiving a codebook comprising a set of visual words, each visual word corresponding to a center of a cluster of local descriptors; calculating, based on a set of user labeled document images, for each visual word of the codebook, a respective frequency distribution of a field position of a specified labeled field with respect to the visual word; loading a document image for extraction of target fields; calculating a statistical predicate of a possible position of a target field in the document image based on the frequency distributions; and detecting, using the trained model, fields in the document image based on the calculated statistical predicate.
Owner:ABBYY DEVELOPMENT INC

System and method for improving reading accuracy and retention

PCT designated stageWO2026136673A1ReadingElectrical appliancesSight wordData science
Computer-based systems and methods for improving reading accuracy and retention for persons with learning disabilities. Articles, documents, stories, novels or other written materials are converted into machine-readable text format by a software program such that customization of the text can be achieved to assist struggling readers in identifying letters and words. At least a portion of the text is displayed using colored lookalike letters and text backgrounds to assist the reader in recognition and identification of letters and words, which will allow the reader to make a color association and allow them to directionally understand which way the letter is facing. Customizable auditory and visual assistance is provided for decoding of high frequency sight words and words for which pronunciation is unknown.
Owner:CLOVER MARTIN LLC

A large model token pruning method based on column-pivot QR decomposition

The application discloses a large model word pruning method based on column-pivoting QR decomposition, and belongs to the technical field of multi-modal large model compression. The application restructures the visual word pruning into a subset selection problem of subspace reservation. Firstly, a double weighting mechanism is designed to fuse the cross-modal correlation guided by text and the intrinsic energy of visual words, to perform adaptive weighting on the original visual feature matrix, and to highlight the semantic core and structural saliency. Secondly, a strategy based on column-pivoting QR decomposition (QRCP) is used to perform greedy rank revealing sorting and extraction on the weighted feature matrix, and a representative word subset that can most approximate the original visual feature space is screened out. The method is plug-and-play and does not need to be retrained, effectively overcomes the feature loss problem of existing methods under a very high pruning rate, greatly reduces the reasoning calculation overhead of a large visual language model, and maintains good model prediction accuracy.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Ultra-high-definition video classification method based on dynamic adaptive masking and word lemma sparsity

The present invention belongs to the technical field of ultra-high-definition video image processing and proposes an ultra-high-definition video classification method based on dynamic adaptive masking and word-element sparsity. The method comprises: using a random mask to randomly mask ultra-high-definition video frames according to a certain ratio, inputting the video frames after random masking into an encoder to obtain intermediate visual features; classifying the intermediate visual features obtained by the encoder to complete fine-tuning of the ultra-high-definition video frames; removing unimportant visual words in the current transformation layer according to decision probabilities and dynamic masks, passing the remaining important visual words to subsequent transformation layers for processing, and completing the classification of a video frame in the ultra-high-definition video; after completing the classification of all frames in a video, directly merging and summarizing the results obtained for each frame individually, and voting to select the class with the most categories as the final ultra-high-definition video classification result. The present invention can reduce the computational complexity of the ultra-high-definition video processing process and ensure computational accuracy.
Owner:SICHUAN NATIONAL INNOVATION VISION UHD VIDEO TECHNOLOGY CO LTD

Systems, methods, and apparatuses for implementing transferable visual words by exploiting the semantics of anatomical patterns for self-supervised learning

Described herein are means for the generation of Transferable Visual Word (TransVW) models through self-supervised learning in the absence of manual labeling, in which the trained TransVW models are then utilized for the processing of medical imaging. For instance, an exemplary system is specially configured to perform self-supervised learning for an AI model in the absence of manually labeled input, by performing the following operations: receiving medical images as input; performing a self-discovery operation of anatomical patterns by building a set of the anatomical patterns from the medical images received at the system, performing a self-classification operation of the anatomical patterns; performing a self-restoration operation of the anatomical patterns within cropped and transformed 2D patches or 3D cubes derived from the medical images received at the system by recovering original anatomical patterns to learn different sets of visual representation; and providing a semantics-enriched pre-trained AI model having a trained encoder-decoder structure with skip connections in between based on the performance of the self-discovery operation, the self-classification operation, and the self-restoration operation. Other related embodiments are disclosed.
Owner:THE ARIZONA BOARD OF REGENTS ON BEHALF OF THE UNIV OF ARIZONA