Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

34 results about "Text query" patented technology

Text Query Applications. The purpose of a text query application is to enable users to find text that contains one or more search terms. The text is usually a collection of documents. A good application can index and search common document formats such as plain text, HTML, XML, or Microsoft Word.

Three-dimensional positioning method, program product, electronic device and storage medium

ActiveCN122115581BPattern recognitionVoxel
The application provides a three-dimensional positioning method, a program product, an electronic device and a storage medium. The method comprises the following steps: obtaining a plurality of RGB image sequences and corresponding depth map sequences of multiple perspectives covering a to-be-positioned region from three-dimensional scene data of the to-be-positioned region; receiving a natural language text query instruction, respectively extracting a text embedding vector of the natural language text query instruction and a visual feature map of the RGB image, and generating a two-dimensional semantic confidence score map of multiple perspectives; calculating the visibility coefficient of a voxel under multiple perspectives by using the depth map sequence, and performing weighted back projection and spatial fusion on the two-dimensional semantic confidence score map by using the visibility coefficient to generate a three-dimensional language evidence body, and analyzing the three-dimensional language evidence body to output a three-dimensional coordinate of a target object in the to-be-positioned region. The zero-sample positioning, fine-grained semantic disambiguation and visibility-aware back projection and evidence fusion mechanism are realized, and the positioning robustness in a cluttered environment is enhanced.
Owner:UNIV OF SCI & TECH OF CHINA

Three-dimensional positioning method, program product, electronic device and storage medium

The application provides a three-dimensional positioning method, a program product, an electronic device and a storage medium. The method comprises the following steps: obtaining a plurality of RGB image sequences and corresponding depth map sequences of multiple perspectives covering a to-be-positioned region from three-dimensional scene data of the to-be-positioned region; receiving a natural language text query instruction, respectively extracting a text embedding vector of the natural language text query instruction and a visual feature map of the RGB image, and generating a two-dimensional semantic confidence score map of multiple perspectives; calculating the visibility coefficient of a voxel under multiple perspectives by using the depth map sequence, and performing weighted back projection and spatial fusion on the two-dimensional semantic confidence score map by using the visibility coefficient to generate a three-dimensional language evidence body, and analyzing the three-dimensional language evidence body to output a three-dimensional coordinate of a target object in the to-be-positioned region. The zero-sample positioning, fine-grained semantic disambiguation and visibility-aware back projection and evidence fusion mechanism are realized, and the positioning robustness in a cluttered environment is enhanced.
Owner:UNIV OF SCI & TECH OF CHINA

Video large language model inference optimization method based on token compression of space-time unification

PendingCN122287875AEngineeringData mining
This application provides a video large language model inference optimization method based on spatiotemporally unified token compression, including: expanding the input video frames into a spatiotemporally unified global visual token set, calculating the saliency weight of each token, selecting tokens to enter a retention pool, and storing the rest in a recycling pool; performing semantic redundancy evaluation and removal on tokens in the retention pool; performing density clustering on tokens in the recycling pool, selecting cluster centers as compensation tokens to re-inject into the retention pool; constructing a text-aware decision score using text queries and tokens in the retention pool, and performing query-guided token merging; inputting the merged tokens into the video large language model for inference to generate the final response. This application solves the performance loss problem of video large language models under ultra-low token retention rates by performing "recycling and refilling" externally and text-aware merging internally. It breaks the assumption of spatiotemporal separation and achieves globally optimal token management by utilizing cross-modal attention and semantic similarity.
Owner:SHANGHAI JIAOTONG UNIV

Information processing system

The present disclosure relates to an information processing system that appropriately manages images captured by a vehicle while reducing costs. The information processing system includes an accumulation unit that accumulates images captured by a camera mounted on a vehicle in the vehicle; a generation unit that generates a feature text indicating a scene feature from image frames included in the images accumulated in the accumulation unit; a feature storage unit that is provided outside the vehicle and stores the feature texts; a query acquisition unit that acquires a search query for searching the images accumulated in the accumulation unit from a user outside the vehicle; a search unit that searches the feature texts stored in the feature storage unit for a feature text that matches the search query; and an output unit that extracts the images corresponding to the feature text that matches the search query from the accumulation unit and outputs the images to the user.
Owner:TOYOTA JIDOSHA KK

Generating multi-order text query results utilizing a context orchestration engine

The present disclosure is directed toward systems, methods, and non-transitory computer-readable media for generating responses to multi-order text queries using a context orchestration engine. For example, the disclosed systems generate context-defining query subcomponents from a multi-order text query, where the context-defining query subcomponents indicate contextual data sources pertaining to their respective portions of the multi-order text query. In addition, the disclosed systems provide or transmit the context-defining query subcomponents to a large language model for domain-specific computer code pertaining to each respective context-defining query subcomponent. The disclosed systems can further execute the generated computer code for each context-defining query subcomponent to access indicated contextual data sources for generating component-specific results. The disclosed systems can also generate a multi-order result to the multi-order text query from the component-specific results.
Owner:DROPBOX INC

Method for retrieving images and eleectronic device performing the method

PendingKR1020260113529AImage queryRadiology
A method performed by an electronic device according to one embodiment may include: an operation of obtaining a text query corresponding to an image query output by a trained first model by inputting an image query into a trained first model; an operation of obtaining one or more images as a first search result for an image query and one or more texts as a second search result for a text query from a database; a set of candidate images determined according to the first search result and the second search result; an operation of obtaining a multimodal score for a set of candidate images output by a trained second model by inputting an image query and a text query into a trained second model; and an operation of determining one or more target images from a set of candidate images based on a multimodal score for a set of candidate images.
Owner:SAMSUNG ELECTRONICS CO LTD

A multi-modal sentiment analysis method and device based on teacher guided distillation and text query selective alignment

ActiveCN122065290BData miningFeature fusion
The application discloses a kind of multi-modal sentiment analysis method and device based on teacher guide distillation and text query selective alignment, it is related to sentiment analysis technical field, comprising: the multi-modal sentiment data is input into trained multi-modal sentiment analysis model, carries out sentiment prediction and outputs sentiment prediction value, wherein, multi-modal sentiment analysis model includes feature extraction module, position encoder, text query selective cross-modal attention module, mode coordination gate module, feature fusion module, time pooling layer and regression prediction head.Based on the above scheme, it has good accuracy and robustness in complex sentiment recognition scene.
Owner:GUANGDONG UNIV OF TECH

An audio text correlation evaluation method based on a multi-expert model, an audio retrieval system, a storage medium and a program product

PendingCN122173674AMetadata audio data retrievalPattern recognitionSemantic alignment
The application belongs to the technical field of audio processing, and a plurality of pre-trained audio text expert models are introduced to construct a multi-expert audio text semantic representation. Meanwhile, the consistency relationship between audio and text at the overall semantic level and the semantic inconsistency between audio and text that is perceptually significant to humans are modeled through a semantic alignment branch and a semantic mismatch branch respectively, and a multi-branch correlation score fusion mechanism is used to output an audio text correlation prediction result that is highly consistent with human subjective evaluation. This method can simultaneously consider semantic consistency and perceptual differences, effectively achieving more accurate and more human subjective evaluation standard-compliant audio text correlation evaluation. Meanwhile, based on the above method, the application constructs an audio retrieval system that evaluates and sorts the correlation between a text query and audio samples in an audio library to realize an audio retrieval function oriented to natural language description.
Owner:HARBIN ENG UNIV

Information processing system

An information processing system includes, a storage configured to store a video captured by a camera mounted on a vehicle within the vehicle, a generator configured to generate a feature text indicating characteristics of a scene from an image frame contained in the video stored in the storage, a feature storage provided outside the vehicle and configured to store the feature text, a query acquirer configured to acquire a search query from a user outside the vehicle to search for the video stored in the storage, a searcher configured to search for the feature text matching the search query from the feature text stored in the feature storage, and an outputter configured to extract the video corresponding to the feature text matching the search query from the storage and output the video to the user.
Owner:TOYOTA JIDOSHA KK

An integrated cross-modal pedestrian retrieval method for detecting and retrieving tasks in parallel

PendingCN122432311AVisual technologyEngineering
The application discloses a kind of detection and retrieval task parallel integrated cross-modal pedestrian retrieval method, it is related to computer vision technical field, solve the technical problem that cross-modal pedestrian retrieval method memory occupation is high, inference time delay is big.The method includes: obtaining the whole image to be retrieved and text query, after processing, input cross-modal weight adapter, generate dynamic weight increment;Convert fixed visual coding into text-guided visual coding;Share the regression query vector of the output of the first layer of coding layer, detect branch predicts target pedestrian bounding box;Based on regression query vector, retrieve branch gets image retrieval feature;Locate and mark target pedestrian in the whole image to be retrieved, output target frame, retrieval score and sorting result;Detection branch and retrieval branch are jointly trained to obtain retrieval model algorithm.The application reduces memory occupation and time delay, meets the demand of pedestrian retrieval.
Owner:UESTC (SHENZHEN) ADVANCED RES INST +1

A data retrieval method and system based on ES large model scoring

The embodiment of the specification provides a data retrieval method and system based on ES large model scoring, comprising: performing vectorization processing on field content in business data that needs to support text query, and storing generated vector data and business data in an Elasticsearch index document; receiving a user query condition, generating a query vector based on the query condition, and constructing an Elasticsearch composite query in combination with a keyword query to obtain a candidate business data set; converting the user query condition into a scoring standard and inputting it into a large model, and scoring data in the candidate business data set according to the scoring standard by the large model; and sorting the candidate business data set according to the scoring result and returning the sorted data. The application can improve the accuracy of the retrieval result and significantly optimize the retrieval experience of the user.
Owner:和创(北京)科技股份有限公司

Video retrieval system and method based on event semantic keyframes and text queries

The present application belongs to the technical field of computer vision or artificial intelligence or intelligent video monitoring, and particularly relates to a video retrieval system and method based on event semantic key frame and text query. The system comprises video preprocessing and event detection module, key frame extraction and semantic annotation module, multi-modal embedding and index construction module connected in sequence; further comprises text query interface, result output module; the method comprises: (1) video preprocessing and event detection; (2) key frame extraction and semantic annotation; (3) multi-modal embedding and index construction; (4) text-driven cross-modal retrieval and result display. The present application realizes rapid, accurate and interpretable positioning from natural language keywords to high-value video evidence without human intervention.
Owner:CHANGFENG DIGITAL TECH (SHANDONG) CO LTD

Bidirectional visual indexing system and method for codebase impact analysis

ActiveUS12675388B1User inputEngineering
Systems and methods for implementing bidirectional visual indexing of a codebase are disclosed. An embodiment of the present invention receives a user input comprising at least one of: a screenshot, a text query, a code diff or a URL link; queries a database to identify one or more related screens based on the user input by comparing vector representations of the user input and a set of data frames; provides a set of metadata for the identified one or more related screens wherein the set of metadata comprises: similarity scores, usage statistics, associated code segments, and network requests; identifies at least one relevant function associated with each of the one or more related screens; performs an impact analysis related to modifying the at least one relevant function using an artificial intelligence (AI) processor; and generates a risk analysis for the modifying of the at least one relevant function.
Owner:MORGAN STANLEY SERVICES GROUP INC

Industrial scene-oriented multi-modal retrieval enhanced model fine-tuning method

The application discloses an industrial scene-oriented multi-modal retrieval enhancement model fine-tuning method, and relates to the technical field of multi-modal large models. The method mainly comprises the following steps: constructing a multi-modal retrieval enhancement model, a multi-modal knowledge base and a fine-tuning data set; obtaining a picture-text instruction from the fine-tuning data set; segmenting a key area of an image and obtaining a high-resolution feature, and constructing a fusion visual feature with a global feature of the image; retrieving a plurality of related text knowledge from the multi-modal knowledge base based on the key area, and constructing a plurality of sets of fusion text features with the text query respectively; inputting the fusion visual feature into a multi-modal language model with the plurality of sets of fusion text features, and updating parameters of a multi-modal retriever according to errors between a plurality of sets of prediction results of the model and standard answers in the case of freezing parameters of the multi-modal language model. The application improves the reasoning efficiency and knowledge utilization capability of the multi-modal model in the industrial scene.
Owner:TIANJIN UNIV OF SCI & TECH

Generative artificial intelligence for content generation with searchable repository

A system may include a content repository that stores multi-modal content comprising text, visual content, and / or audio content. The system may include a processor programmed to: access a prompt comprising a text query, a visual query comprising text that describes a visual to be found, and / or an audio query comprising text that describes audio to be found, execute a language model based on the prompt to identify content from the content repository, receive, from the language model, a request for a callback function that seeks additional information to satisfy the multi-modal query, execute the callback function to obtain the additional information and provide the additional information to the language model in response to the request for the callback function, re-execute the language model based on the multi-modal query and the additional information, obtain, from the language model, content responsive to the prompt based on the additional information.
Owner:REVE AI INC

Deeply guided first-person hand-object interaction understanding method

PendingCN122453904ALinguistic modelRgb image
The application discloses a depth-guided first-view hand-object interaction understanding method, and relates to the fields of computer vision and multi-modal artificial intelligence. A first-view RGB image, a monocular depth map aligned with the first-view RGB image and a text query instruction are acquired; feature encoding is respectively performed on the RGB image and the depth map to obtain RGB visual tokens and depth tokens; context modeling is performed on the RGB visual tokens to obtain global semantic features; proxy tokens are generated based on the depth tokens to extract geometric information, and local fusion of RGB features and depth information is realized through a proxy attention mechanism; global modulation is performed on the fused visual features by using depth information prediction feature modulation parameters; and the fused visual tokens and the text tokens are input into a visual language model for reasoning generation to obtain interaction semantic descriptions and spatial positioning results. The application can introduce depth geometric information while maintaining the semantic ability of the visual language model, and improve the target positioning and semantic understanding ability in the first-view interaction understanding task.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

A semantic picture editing method based on diffusion model

The application discloses a semantic picture editing method based on a diffusion model, belongs to the technical field of computer vision and artificial intelligence, and comprises the following steps: constructing a text encoder and an image encoder, embedding labels into a vector space through the text encoder, calculating through the image encoder and the labels, and generating a picture mask; constructing a denoising diffusion probability model, encoding an input image based on the denoising diffusion probability model, and obtaining latent features corresponding to the input image; performing decoding on the denoising diffusion probability model based on the picture mask and the latent features corresponding to the input image, obtaining an inferred mask; and replacing the background of the input image with pixel values in the encoding process based on the inferred mask and mapping the pixel values back to original pixels to obtain a new picture. The application discloses a semantic-guided picture editing method based on a diffusion model. According to the method, which regions of an input image should be edited can be automatically found according to a text query.
Owner:HEFEI UNIV OF TECH

Data processing method and system for metabolic cabin skeletal point action recognition

PendingCN122347689AStreaming dataData set
This application provides a data processing method and system for skeletal point action recognition in a metabolic chamber, relating to the field of behavior recognition. The method includes: processing RGB video stream data, determining atomic state feature vectors and object feature vectors, inferring semantic-level action category labels and spatiotemporally aligning them with 3D skeletal target point data to generate a training dataset; performing fully supervised training on a neural network model to obtain a target skeletal perception network; inputting real-time 3D skeletal target point data into the target skeletal perception network and outputting action categories; converting action categories within a continuous time period into structured text behavior logs; and performing a matching query to return the corresponding text behavior logs when a user's text query command is received. This application eliminates the risk of visual privacy leakage, reduces the computational and storage requirements for edge hardware, accurately distinguishes complex actions that are geometrically similar but involve different interactive objects, and achieves millisecond-level fast backtracking.
Owner:ANHUI HONGYUAN JUKANG MEDICAL TECH CO LTD

Search enhancement generation methods, apparatus, equipment, media and program products

PendingCN122285830AText databaseUser input
This application provides a retrieval enhancement generation method, apparatus, device, medium, and program product, relating to the field of artificial intelligence technology, for improving the quality of generated answers. The specific technical solution is as follows: A user-inputted query question is input into an answer model; the answer model obtains modal requirement information corresponding to the user-inputted query question; the answer model performs a text query in a text database based on the query question to obtain at least one response text fragment; the answer model performs contextual reconstruction processing on the at least one response text fragment based on a target reconstruction strategy matching the modal requirement information to obtain a target answer that conforms to the user's modal requirement information. This application is applied to intelligent search scenarios.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

Data query method and device, electronic equipment and storage medium

The application relates to the technical field of big data, and provides a data query method and device, electronic equipment and a storage medium, which comprise the following steps: in response to a data query request, a target connected graph is constructed based on a plurality of original vector data included in different index objects that have been written, the original vector data comprises a query text vector and a corresponding query result vector, the data query request carries a target query text vector and a query result quantity requirement; based on the target connected graph, a plurality of candidate coordinate positions adjacent to the coordinate position of the target query text are found, and at least one target coordinate position is determined from the plurality of candidate coordinate positions based on the query result quantity requirement; and the query result vector of the query text vector corresponding to the at least one target coordinate position is determined as the query result of the data query request. The application combines the near neighbor text query and the graph algorithm, reduces the time complexity of data query, and improves the data query efficiency.
Owner:MIDEA NETWORK INFORMATION SERVICE (SHENZHEN) CO LTD

An enhanced shallow KV Cache compression method and system based on temporary storage screening and residual compensation

The application relates to an enhanced shallow KV Cache compression method and system based on temporary storage screening and residual compensation. A client (such as a user terminal device) is responsible for acquiring user inputted long text query data and sending the long text query data to a server end. The server end is provided with a large language model with a Transformer architecture. The application is executed in a computing device of the server end. Through differential and dynamic compression scheduling of shallow and deep KV Cache (key value cache) in an LLM inference process, high throughput and high precision long text inference can be completed under the condition of limited physical display memory of hardware, and finally, the generated text result is returned to the client.
Owner:SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN +1

Apparatus and method for video moment retrieval and highlight detection using inverted token augmentation

A method and a device for video moment retrieval and highlight detection may include obtaining an inverted token of a text query based on an original token of the text query, obtaining video feature vectors of video clips, and performing an interaction between the text query and video clips by simultaneously using the original token, the inverted token and the video feature vectors.
Owner:ELECTRONICS & TELECOMM RES INST

A multi-modal information processing method based on knowledge retrieval enhancement

The application discloses a kind of multi-modal information processing methods based on knowledge retrieval enhancement, belong to multi-modal information processing technical field. For the single knowledge retrieval path of prior art, the problems of insufficient environment and entity modeling and lack of reliability control, the application builds multi-modal cultural knowledge base containing environment node, entity node and hierarchical knowledge fragment, establishes the association between nodes;Based on the semantic features of text Query and the local entity retrieval value of image, a routing signal is generated, and an entity direct search or global context indirect search strategy is adaptively selected to obtain candidate knowledge;Purify knowledge, use cross-modal consistency check to filter visual verifiable knowledge, and perform controlled retention on non-visual extended knowledge, finally input the purified knowledge set together with the input into the large model reasoning to obtain the reasoning result. The application improves the accuracy and reliability of knowledge positioning in complex scenarios. It is suitable for cultural scene analysis, travel and travel question and answer and intelligent reasoning field.
Owner:HARBIN INST OF TECH

A dialogue text normative analysis method and device and a storage medium

The application relates to a dialogue text normative analysis method and device and a storage medium, and is applied to the technical field of text analysis, and comprises the following steps: storing dialogue texts in an Elasticsearch cluster according to speaking roles, the Elasticsearch cluster is based on Lunce, supports self-defined word segmentation and query plug-ins, on the basis, the dialogue texts are stored according to the roles, the problem that a word segmentation plug-in cannot distinguish and store role texts based on the full-text retrieval technology Lunce in the prior art is solved, and the application sets self-defined system reserved words, sets a query syntax according to the self-defined system reserved words, generates a syntax tree through Antlr4 technology, obtains a query statement that can be recognized by the Elasticsearch cluster through the syntax tree, compared with the construction of the query syntax in the prior art which adopts a common tree type data structure, the construction of the dialogue text query syntax in the application adopts the Antlr4 technology which is commonly used in the industry, and storage, checking and conversion are relatively simple.
Owner:SHANGHAI ZHONGTONGJI NETWORK TECH CO LTD