Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

17 results about "Visual Word" patented technology

Visual words, as used in image retrieval systems, refer to small parts of an image which carry some kind of information related to the features (such as the color, shape or texture), or changes occurring in the pixels such as the filtering, low-level feature descriptors (SIFT, SURF, ...etc.).

Method and device for positioning and mapping industrial automobile crane based on multi-sensing fusion

The invention discloses an industrial automobile crane positioning and mapping method and device based on multi-sensing fusion, and relates to the field of machine sensing, and the method comprises the following steps: initializing a pose and offset of an inertial measurement unit by using an input laser point cloud and a pre-integration result of the inertial measurement unit; correcting laser point cloud motion distortion by using an inertial measurement unit, extracting a dynamic object point cloud, and synchronously extracting laser features and visual features with depth; calculating a key frame pose in real time by minimizing a visual re-projection error and an inertial measurement unit pre-integration error, and performing frame-to-map matching by taking the key frame pose as an initial value to further optimize the pose; a visual word bag model is utilized to quickly retrieve a closed-loop candidate frame, after laser passes geometric matching to verify a closed loop, a laser radar factor, a visual factor, an inertial measurement unit factor and a closed-loop factor are fused, a global factor graph is constructed and optimized, and a pose and point cloud map is updated. According to the method, the environment sensing capability of the industrial automobile crane can be improved, and the sensing accuracy and robustness are improved.
Owner:NANKAI UNIV +1

Training method of image encoder, image processing method, electronic device, storage medium and program product

The embodiment of the application provides a kind of training method of image encoder, image processing method, electronic equipment, storage medium and program product, it is related to image processing technical field, the training method of image encoder includes: training image is input image encoder, visual word symbol for representing image feature is extracted by the image encoder;Clothing text word symbol is generated for training image;The clothing text word symbol is used to describe the clothing feature of the person in the training image;The model parameter of the image encoder is iteratively trained using a combination loss function;The combination loss function is used to make the visual word symbol of the same person of multiple training images be smaller, and make the visual word symbol and the clothing text word symbol of the same training image be larger difference.It reduces the degree of dependence on visual information by introducing text modal, increases the precision of image encoder to identify the same person under different clothing, improves the precision and robustness of person re-identification.
Owner:PEKING UNIV

Electric power vision large model multi-scale semi-supervised target detection method and system

The invention discloses an electric power vision large model multi-scale semi-supervised target detection method and system, and the method comprises the steps: training a multi-scale vision word segmentation device based on a vector quantization knowledge distillation framework, and constructing an electric power semantic enhanced multi-scale vision codebook, the multi-scale visual codebook comprises a top-layer codebook used for encoding a global structure of the power equipment and a bottom-layer codebook used for encoding defect local features of the equipment; based on the multi-scale visual codebook, adopting a mask image modeling task to pre-train a visual large model on an unmarked power image; performing semi-supervised fine tuning on the visual large model by using the labeled electric power image data and the unlabeled electric power image data; using an SMLS fine tuning strategy to update the visual large model parameters; and inputting a to-be-detected power image into the updated visual large model, and outputting a detection result of the multi-scale target. According to the method, the problems of scarcity of annotation data, model semantic gaps and high calculation complexity in an electric power scene in the prior art are solved.
Owner:STATE GRID ELECTRIC POWER RES INST +1

Training methods for scene classification models, remote sensing image scene classification methods and equipment

This invention provides a training method for a scene classification model, a remote sensing image scene classification method, and an apparatus. The specific implementation includes: acquiring multiple first image blocks and their respective sample labels; using a feature extraction network to extract features from each first image block to obtain local features; generating a visual word histogram for each first image block based on its local features; using a first classification network and a second classification network based on the visual word histogram to obtain a first classification result and a second classification result corresponding to each scene to be classified; determining a first loss value for the first classification result and a second loss value for the second classification result based on the correlation between the first classification result, the second classification result, the sample labels, and the multiple scenes to be classified; and adjusting the parameters of the first classification network and the second classification network according to the first loss value and the second loss value, respectively.
Owner:BEIJING DATA INTELLIGENCE INFORMATION TECH CO LTD

Method for hallucination detection and suppression of multi-modal large model based on attention time difference

The application discloses a hallucination detection and suppression method based on attention timing difference of a multimodal large model, which comprises the following steps: inputting text word elements of a prompt word and visual word elements of an image into a multimodal large model to obtain an internal attention graph in a decoding stage of the multimodal large model; then, calculating an attention proportion of the visual word elements in a current generation moment and performing difference with the proportion in a previous moment; if the difference exceeds a set threshold, it is determined that the visual word elements are related; when the visual related word elements are identified, a second forward propagation with visual enhancement is performed to obtain more accurate output; if not, the next generation is directly entered. According to the application, the visual related word elements in text generation are identified and refined through attention timing difference, twice forward propagation is adopted, the visual attention of the second forward propagation is enhanced based on the visual attention graph of the first forward propagation, and hallucination can be identified and corrected without reducing the language expression ability.
Owner:HANGZHOU DIANZI UNIV

A semantic selection and quota control method and system based on user vocabulary state

PendingCN122311131AData miningTarget text
This invention discloses a semantic selection and quota control method and system based on user vocabulary state. The method includes: acquiring target text and identifying word-level candidate sequences in the text, where the candidate sequences preserve the structural positional relationships of words within the text; constructing a user vocabulary state space, modeling the vocabulary selection process as a path construction process under the constraints of the state space Ω, where candidate words serve as path nodes, and whether a node enters the path depends on its state label and constraint function; dividing the candidate sequences based on the state space Ω to obtain a set of known words and a set of unknown word candidates; during the path construction process, imposing constraint functions on the unknown word candidate set, and selecting words from it to generate a new word set; merging the new word set with the known word set to generate a visual word path, and generating semantically enhanced text data for computer display or output based on the visual word path.
Owner:CHUANGZHI YUNWEI (BEIJING) TECHNOLOGY CO LTD

A method and related equipment for semantic understanding of long-term traffic videos

This application provides a method and related equipment for semantic understanding of long-term traffic videos, relating to the fields of intelligent transportation and computer vision technology. This application reconstructs a target long-term traffic video sequence into a target sparse traffic visual word sequence involving dynamic traffic visual concepts. Then, it performs cross-timescale sparse attention feature learning on the target sparse traffic visual word sequence across multiple traffic state perception dimensions. Based on the assigned desired video semantic understanding task, it fuses the local attention output feature sequences of each of the learned traffic state perception dimensions to obtain the target traffic video semantic feature sequence, minimizing redundant information interference during the traffic video semantic understanding process. Finally, it invokes a large language model to perform cross-modal attention interaction reasoning based on the target traffic video semantic feature sequence and the desired video semantic understanding task, effectively balancing the efficiency and accuracy of video semantic understanding.
Owner:BEIHANG UNIV

A core subgraph-based training-free multi-modal large language model fine-grained positioning method

PendingCN122451086ALinguistic modelAlgorithm
The application belongs to the technical field of computers, and particularly relates to a core subgraph-based training-free multi-modal large language model fine-grained positioning method. The method comprises the following steps: 1) using a pre-trained multi-modal large language model to respectively perform feature coding on input images and texts, and respectively dividing the images and texts into a plurality of word units to construct visual and text word unit vector sequence sets; 2) constructing a visual semantic graph structure according to the semantic similarity relationship between visual word units, taking each visual word unit as a graph node, and taking the similarity between the nodes as an edge weight, so as to form a normalized adjacency matrix; 3) calculating the importance score of each visual word unit, and taking a region specified by a user in an arbitrary shape as a core subgraph; 4) performing a visual attention diffusion process in the visual semantic relationship graph based on the core subgraph, and determining a subgraph according to the importance distribution after diffusion; and 5) inputting the screened subgraph and the coded text word unit into a large language model to generate an answer through decoding.
Owner:NANKAI UNIV

SLAM loopback detection method and system based on pipeline topology constraint

The invention discloses a pipeline topology constraint-based SLAM loopback detection method and system, and belongs to the technical field of robot synchronous positioning and mapping. Comprising the following steps: acquiring a prior topological map of a pipeline scene; retrieving loopback candidate frames through a visual word bag model; the method is characterized in that SLAM estimation poses of a current frame and a candidate frame are mapped to nodes of a topological map, and a topological path distance between the two nodes is calculated; and verifying and screening the loopback candidates based on a consistency comparison result of the distance and a preset threshold value. The device is mainly used for providing high-reliability loopback detection capacity for the oil and gas pipeline inspection robot. According to the method, the prior topology constraint is introduced as a strong filter, false loops can be effectively eliminated from the source, and the robustness and the positioning precision of the SLAM system in a structured industrial environment are remarkably improved.
Owner:SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI +1

Information retrieval method, device, storage medium, and program product

PendingCN122332597AData setData mining
This invention provides an information retrieval method, device, storage medium, and program product. The method includes: acquiring a representation model trained on a training dataset, wherein different retrieval tasks correspond to different data modality combinations in the training data pairs, and the text data used to describe image data in the image-text pairs includes visual words that meet a set density condition; acquiring retrieval input data and multiple retrieved data corresponding to the target retrieval task; determining the embedding representation vectors corresponding to the retrieval input data and the multiple retrieved data respectively through the representation model; and selecting the target retrieved data from the multiple retrieved data based on the similarity between the two. The training dataset uses dense multimodal knowledge and training data from various retrieval tasks, improving the multimodal alignment capability of the representation model, enabling it to adapt to different types of retrieval tasks, and improving the accuracy of retrieval results.
Owner:ALIBABA CLOUD COMPUTING CO LTD

A new method for open-vocabulary semantic segmentation of remote sensing images

The application discloses a new method for remote sensing image open vocabulary semantic segmentation. The method comprises the following steps: a dense feature calibration stage, which improves the spatial reliability of low-resolution visual words by abnormal feature repair and semantic correlation attention mechanism; an arbitrary resolution feature up-sampling stage, which reconstructs low-resolution features into high-resolution dense representations by image-guided localized cross-attention guided by input image structure clues; and a coordination aggregation stage, which aggregates the features extracted from multiple geometrically transformed views after inverse transformation alignment to suppress perspective-related fluctuations. The application does not require remote sensing pixel labeling, can be used immediately after migration among the CLIP series backbone, and is suitable for open vocabulary pixel-level recognition in multiple types of remote sensing scenes such as satellite images and unmanned aerial vehicle aerial images.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Synchronous positioning and mapping method and system for multi-source information fusion

The invention provides a multi-source information fusion SLAM (Simultaneous Localization and Mapping) algorithm and system. Vision, laser radar and IMU (Inertial Measurement Unit) data are deeply fused. A system framework comprises data preprocessing, pose estimation, back-end optimization and a hardware platform. The data module preprocesses the laser point cloud; an improved ICP laser odometer and a visual odometer (combining optical flow tracking and angular point detection) are adopted for pose estimation, and error state Kalman filtering fusion is carried out. And back-end optimization is based on laser point cloud and visual word bag model design loopback detection, accumulative errors are reduced, and a color point cloud map is generated. A hardware platform is carried on an unmanned aerial vehicle and an unmanned vehicle, a multi-sensor scanner is constructed, sensor calibration is carried out, and remote networking and ROS distributed configuration are adopted. The method is high in positioning precision, high in robustness and centimeter-level in mapping precision in a complex environment, information of multiple sensors is effectively fused, the generated color RGB point cloud map is close to a real space, construction of a digital twinborn model is facilitated, and AR / VR technology development is promoted.
Owner:项朝阳

System and method for improving reading accuracy and retention

PCT designated stageWO2026136673A1ReadingElectrical appliancesSight wordData science
Computer-based systems and methods for improving reading accuracy and retention for persons with learning disabilities. Articles, documents, stories, novels or other written materials are converted into machine-readable text format by a software program such that customization of the text can be achieved to assist struggling readers in identifying letters and words. At least a portion of the text is displayed using colored lookalike letters and text backgrounds to assist the reader in recognition and identification of letters and words, which will allow the reader to make a color association and allow them to directionally understand which way the letter is facing. Customizable auditory and visual assistance is provided for decoding of high frequency sight words and words for which pronunciation is unknown.
Owner:CLOVER MARTIN LLC

A large model token pruning method based on column-pivot QR decomposition

The application discloses a large model word pruning method based on column-pivoting QR decomposition, and belongs to the technical field of multi-modal large model compression. The application restructures the visual word pruning into a subset selection problem of subspace reservation. Firstly, a double weighting mechanism is designed to fuse the cross-modal correlation guided by text and the intrinsic energy of visual words, to perform adaptive weighting on the original visual feature matrix, and to highlight the semantic core and structural saliency. Secondly, a strategy based on column-pivoting QR decomposition (QRCP) is used to perform greedy rank revealing sorting and extraction on the weighted feature matrix, and a representative word subset that can most approximate the original visual feature space is screened out. The method is plug-and-play and does not need to be retrained, effectively overcomes the feature loss problem of existing methods under a very high pruning rate, greatly reduces the reasoning calculation overhead of a large visual language model, and maintains good model prediction accuracy.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Data processing method and device, storage medium and electronic equipment

The invention discloses a data processing method and device, a storage medium and electronic equipment, and relates to the technical field of recommendation artificial intelligence. The method comprises the following steps: fusing spatial layout information and context image semantic information of a plurality of full-slice images through a three-dimensional position coding and vector fusion mechanism to obtain a fusion vector; dynamically aggregating the visual lexical elements formed by the fusion vectors into a compressed representation with a fixed length through a learnable lexical element compression mechanism; and inputting the compressed representation into a multi-modal large model to obtain a structured analysis report. According to the method, the technical problems that the calculation amount is too large, spatial information fusion is insufficient, key information is easy to lose, and report accuracy and reliability are poor when a full-slice image is processed in an existing mode can be solved. Therefore, the diagnosis report generated by using the method can meet clinical precise diagnosis requirements.
Owner:CHINA MOBILE COMM LTD RES INST +1

Visual word segmentation device training method, device, equipment, medium and program product

The invention discloses a training method and device of a visual word segmentation device, equipment, a medium and a program product, and relates to the field of artificial intelligence. The method comprises the following steps: acquiring a sample image block in a sample image; encoding the sample visual features of the sample image blocks; obtaining a quantization index vector of the sample image block based on the sample visual feature and at least two coding features in the codebook; the quantization index vector is related to the probability that each coding feature of the at least two coding features serves as the coding feature closest to the sample visual feature; obtaining sample quantization features of the sample image blocks based on the quantization index vectors; the sample quantization feature is obtained by fusing at least two coding features; decoding the sample quantization feature to obtain a predicted image block; and training model parameters and at least two coding features of the visual word segmentation device by taking reduction of the difference between the prediction image block and the sample image block as a training target.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Multi-modal data processing method and electronic equipment

The invention provides a multi-modal data processing method and electronic equipment. The multi-modal data processing method comprises the steps of obtaining a to-be-processed image and task indication information; the task indication information represents that a target task is executed at least based on the to-be-processed image; segmenting the to-be-processed image into a plurality of image blocks, and determining a plurality of visual lexical elements corresponding to each image block and probability distribution corresponding to the plurality of visual lexical elements; the probability corresponding to the visual lexical element represents the probability that each image block belongs to the visual lexical element; generating a visual embedding feature corresponding to each image block based on the feature corresponding to each visual lexical element in the plurality of visual lexical elements and the probability distribution, and obtaining a plurality of visual embedding features of the plurality of image blocks; and processing the plurality of visual embedded features and the textual features of the task indication information by using a multi-modal model to generate an execution result of the target task.
Owner:LENOVO (BEIJING) LTD