Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

203 results about "Text query" patented technology

Text Query Applications. The purpose of a text query application is to enable users to find text that contains one or more search terms. The text is usually a collection of documents. A good application can index and search common document formats such as plain text, HTML, XML, or Microsoft Word.

Multi-modal geographic positioning method based on knowledge graph

The invention provides a multi-modal geographic positioning method based on a knowledge graph. The method comprises the following steps: graph construction: constructing a geographic knowledge graph; image positioning: calculating the entity similarity between the to-be-queried image and each sub-image so as to select an image positioning candidate sub-image, calculating the image matching similarity between the to-be-queried image and each image positioning candidate sub-image so as to determine an image positioning target sub-image, and taking the latitude and longitude coordinates of each image positioning target sub-image as a positioning result; text positioning: converting a to-be-queried text into map representation, calculating the map similarity between a text query sub-graph and each sub-graph to select a text positioning candidate sub-graph, and calculating the semantic similarity between the to-be-queried text and each text positioning candidate sub-graph, and calculating a text matching similarity corresponding to each text positioning candidate sub-graph based on the graph similarity and the semantic similarity corresponding to each text positioning candidate sub-graph so as to determine a text positioning target sub-graph, and taking latitude and longitude coordinates of each text positioning target sub-graph as a positioning result.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Method and system for constructing vector databases used for converting free text queries to cyber language queries

A system and method for querying data sources for cybersecurity analysis is presented. The system and method include: receiving security logs from at least one data source, wherein the security logs lack pre-defined schema; generating a schema of the security logs based on at least a type of data of the security logs, wherein the generated schema includes fields of the security logs and values of the fields; embedding field vectors, wherein each field vector is a vector representation of a value of each respective field; embedding value vectors, wherein each value vector is a vector representation of a natural language description of each value in each respective field; and generating a query in a cyber language query, using an AI system, for execution on at least one target data source based, in part, on the generated schema, the embedded field vectors, and the embedded value vectors.
Owner:VEGA CYBER SOLUTIONS LTD

Weak supervision online video moment positioning method and system based on memory perception

The invention relates to a weak supervision online video moment positioning method and system based on memory perception, and belongs to the technical field of artificial intelligence, and the method comprises the steps: carrying out the multi-modal feature fusion of a given video and a text query thereof, and obtaining the unified frame level representation at each stage; inputting the fused features into an offline module and an online module in an integral and frame-by-frame manner by using an offline guide online model architecture; in the off-line module, generating a Gaussian mask to reconstruct query of a covered part of words, and obtaining a proposal of an action starting moment; in the on-line module, the long-term historical memory in the window is used for enhancing the score, the attention weight of the score in the window is dynamically generated, and the score of the current frame is calculated in a weighted mode; taking the proposal obtained by the offline module as a pseudo tag, and providing supervision information for the score sequence of the online module; and high-performance weak supervision on-line moment positioning can be completed only by independently deducing the on-line module. The expansion capability and the application value of the model are remarkably improved.
Owner:SHANDONG UNIV

Distributed Hybrid Search for Language-Agnostic, Real-Time Information Retrieval

A computer-implemented method for performing searches in a document database is disclosed. The method comprises automatically detecting a line of business associated with a user, receiving a text query from the user, and generating a query embedding from the text query. The method further comprises scoring entries in a reverse index using a hybrid scoring function. The reverse index comprises titles, title embeddings, sentences, sentence embeddings, and entity tags corresponding to documents in the document database. The hybrid scoring function is used to generate a score based both on a keyword match score between the text query and the reverse index and on a cosine similarity score calculated from embeddings in the text query and in the reverse index. The method also comprises ranking scores for sentences in the document database, and displaying a sentence associated with a top score to the user.
Owner:DELL PROD LP

Electronic device for at least one of video moment retrieval and highlight detection and operation method thereof

Proposed is an electronic device for at least one of video moment retrieval and highlight detection which includes a storage unit and a processor, wherein the processor obtains a plurality of first video features from a video, obtains a text query feature from a text query, obtains a plurality of weights from the plurality of first video features and the text query feature, obtains a plurality of second video features from the plurality of weights and the plurality of first video features, obtains a plurality of third video features from the plurality of second video features by using an encoder, obtains a plurality of fourth video features from the plurality of third video features and a time query by using a decoder, and selects at least one of time points or time periods of the video by using the plurality of fourth video features.
Owner:RES & BUSINESS FOUND SUNGKYUNKWAN UNIV +1

Multimodal context selection for large language model based resolutions addressing technical issues

A method for technical issue resolution. The method includes: receiving, from a user, a text query concerning a technical issue; obtaining query-related context relevant to the text query; and processing, through a large language model (LLM), the text query and the query-related context to produce a multimodal query response used by the user to address the technical issue. More specifically, embodiments described herein utilize text topic and zero shot classification models to translate multimodal technical documentation (e.g., including text and images) into topic relevant metadata; and process queries, pertaining to technical issues, using a multimodal LLM provided with query-related text and image context derived from said topic relevant metadata.
Owner:DELL PROD LP

Visual positioning method based on semantic comprehension and attribute distinguishing enhancement

The invention belongs to the technical field of visual positioning, and relates to a visual positioning method based on semantic comprehension and attribute distinguishing enhancement. The framework mainly comprises a feature coding module, a semantic sensitive data enhancement module, a fine-grained attribute guiding module and a multi-stage cross-modal decoder module. Specifically, the feature coding module performs feature coding on an input graph. The semantic sensitive data enhancement module generates a plurality of queries which are consistent with long text semantics by keeping the consistency of spatial relation words in combination with a large language model, so that a training data set for the long text is expanded. The fine-grained attribute guiding module extracts attribute prior information from a text query in combination with a text graph model and an image encoder, constructs a visual feature representation with higher discrimination by using the information guiding model, and generates a target query with attribute difference at the same time.
Owner:DALIAN UNIV OF TECH

Target-level three-dimensional point cloud cross-modal semantic retrieval method and device and electronic equipment

The invention discloses a target-level three-dimensional point cloud cross-modal semantic retrieval method and device and electronic equipment. The target-level three-dimensional point cloud cross-modal semantic retrieval method comprises the steps that independent three-dimensional target point clouds are segmented from three-dimensional scene point clouds, and unique identifiers are given to the independent three-dimensional target point clouds; projecting each target point cloud to three orthogonal two-dimensional observation planes to generate a multi-view composite image; a pre-trained vision-language basic model is adopted to code and fuse the synthesized image, and a unified multi-modal feature vector is generated; a vector database in which the feature vectors are associated with their identifiers is constructed. And encoding a natural language query text into a text query vector by using the model, calculating the semantic similarity between the text query vector and the feature vector in the database, and positioning and returning the corresponding three-dimensional target point cloud. According to the method, accurate and efficient cross-modal retrieval from a natural language to a three-dimensional point cloud target is realized, the problem that a traditional method is difficult to support semantic fine-grained retrieval of the three-dimensional target point cloud is solved, and a key technical support is provided for training data management in the fields of intelligence and the like.
Owner:MOLAR INTELLIGENCE INFORMATION TECHNOLOGY (HANGZHOU) CO LTD

Three-dimensional scene local editing method based on text control

A three-dimensional scene local editing method based on text control comprises the steps that a three-dimensional Gaussian semantic field is constructed based on text query, and accurate three-dimensional object masks are directly generated by embedding semantic features into anisotropic Gaussian primitives and combining multi-view semantic supervision and cosine similarity calculation of a CLIP model; screening an effective view angle by taking a target object as a center, determining an effective view angle range through azimuth clustering analysis and kernel density estimation, and generating a continuous dense surround-view camera track; an image grid technology is adopted to organize and edit visual angles, a submerged space alignment module GLAM fusing global-local attention is utilized to synchronously process multiple frames of images, style benchmarks are unified and geometric consistency is enhanced through a cross-visual-angle attention mechanism, a multi-visual-angle consistent editing result is generated, and only Gaussian parameters corresponding to a target object are finely adjusted, so that the target object can be edited more accurately. And the background area is kept unchanged. According to the method, the boundary definition of local editing is remarkably improved, and efficient and accurate three-dimensional content modification of a multi-target complex scene is achieved.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL +1

Dynamic document annotation system

A set of locations in a data object to be annotated is identified as corresponding to metadata of the data object. A natural language text query is generated using the metadata of a data object. A set of scores is generated for the set of locations using a generative neural network, and the set of scores indicate whether individual candidate locations satisfy the natural language text query. Based on the set of scores, a location in the data object is annotated to generate an annotated location as corresponding to the metadata.
Owner:CITIGROUP

Systems and methods for orchestrating interaction with an artificial intelligence application

Systems and methods for orchestrating interaction with an artificial intelligence (AI) application in a contact center environment receive, via an AI agent, a voice message from a user; convert the message from voice to text; generate an initial computational inference process based on the text message; determine whether or not all information required to execute the initial computational inference process is available to the processor; when a determination is made that all information required is available: execute the initial computational inference process; generate a text reply based on the initial computational inference process; convert the text reply to a voice reply; and send the voice reply to the user via the AI agent; when a determination is made that information is unavailable: generate a text query requesting the information; convert the text query to a voice query; and send the voice query to the user via the AI agent.
Owner:THE BANK OF NEW YORK MELLON

Image text alignment method based on multi-modal large language model

PendingCN121542772AText alignmentFeature vector
The invention relates to the technical field of image texts, in particular to an image text alignment method based on a multi-modal large language model, which comprises the following steps of: segmenting an image into local areas and encoding the local areas into visual feature vectors by adopting a pre-trained visual encoder; visual and text features are mapped to the same dimension space through a learning projection layer, an InfoNCE loss function is adopted to maximize the similarity of a positive sample pair, dynamic interaction between text query and an image area is achieved through a cross attention mechanism, language module parameters are frozen to pre-train a visual encoder, an overall module is combined to be fine-tuned, and text query is completed. An error case is analyzed, and a verified and improved alignment module is obtained through a data enhancement optimization module. The method solves the problems that a traditional image text alignment method is incomplete in manual feature extraction and low in precision, an existing deep learning model is limited in cross-modal semantic understanding ability, and the generalization ability is poor when facing long-tail data.
Owner:GUANGDONG UNIVERSITY OF FOREIGN STUDIES

Information processing

To-be-queried information is received by a server device. The to-be-queried information is transmitted from a terminal device. A multimodal information search is performed by the server device based on the to-be-queried information to obtain multimodal search results. A content digest extraction of a target text in the multimodal search results is performed to obtain one or more content digest fragments. Based on the one or more content digest fragments, a text query result is generated. When a first preset number of search results in the multimodal search results lack a rich media query result, key information extraction on the text query result is performed to obtain key description information. A supplemental rich media query result is acquired according to the key description information. The text query result and the supplemental rich media query result are fused into a target query result that is transmitted to the terminal device.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Generating video descriptions using a machine learning model

The present disclosure describes techniques for generating video descriptions using a machine learning model. A plurality of sets of visual tokens corresponding to a plurality of frames of a video is generated. A first type of tokens is generated by implementing temporal pooling on the plurality of sets of visual tokens corresponding to the plurality of frames. A second type of tokens is generated by compressing each of the plurality of sets of visual tokens corresponding to each of the plurality of frames. A third type of tokens is generated by applying cross-attention between each of the plurality of sets of visual tokens and a fourth type of tokens including text tokens generated based on an input text query. A text description of the video is generated based on the first type of tokens, the second type of tokens, the third type of tokens, and the fourth type of tokens.
Owner:LEMON INC(GB)

Computer-implemented method and system for behavior planning of a vehicle

A concept for behavior planning of a vehicle in a traffic scene is proposed that extends the planning capabilities of a rule-based planner so that the rule-based planner can also be used for behavior planning in more complex scenarios. A rule-based planner (50) is used which, on the basis of a scene representation (2) of the traffic scene (1), carries out a behavioral planning (6) which simulates at least one planning scenario from a planner-specific set of predetermined planning scenarios, According to the invention, a text-based neural network (30) is used to select the at least one planning scenario. For this purpose, at least one text query (3) describing the traffic scene is generated for the neural network (30) based on the scene representation (2), with which a behavioral recommendation (4) is requested that is supported by the planner-specific set of predefined planning scenarios. Based on the text query (3), the neural network (30) generates a text-based behavioral recommendation (4) that is assigned to at least one of the predefined planning scenarios. The behavioral planning (6) of the rule-based planner (50) is then based on the planning scenario thus selected.
Owner:ROBERT BOSCH GMBH

System for realizing intelligent office guidance of office staff based on government affair element universe digital person

The invention provides a government affair element universe digital person-based system for realizing intelligent guide of affair handling personnel, and belongs to the technical field of government affair service intellectualization. The system deeply integrates element universe, intelligent interaction, big data and artificial intelligence technologies, realizes natural interaction with the affair handling personnel by means of an intelligent virtual digital person, and realizes intelligent guide of the affair handling personnel. And the government affair service process can be smoothly completed. A local intelligent knowledge base system and a knowledge base exclusive to multiple government affair departments are built, and a DeepSeek large model is applied. Voice or text query of business knowledge, handling processes, policies and regulations is supported, and contents cover latest policies of tax payment services, meta universe, offline business handling places and modes and the like. The digital human has abundant functions, provides business consultation, hall navigation guidance and intelligent explanation, can also perform daily question and answer interaction, and creates a relaxed communication atmosphere. And after logging in the meta universe cloud hall, the accompanying real-time efficient interactive guide service can be obtained by clicking the guide digital person, so that the government affair service efficiency and the user satisfaction are remarkably improved.
Owner:SHANDONG INSPUR COMML SYST CO LTD +1

Multi-modal retrieval enhancement generation method and system based on multi-expert model

The invention discloses a multi-modal retrieval enhancement generation method and system based on a multi-expert model. The multi-modal retrieval enhancement generation method comprises the following steps: S1, processing input multi-modal data by using a visual language model and an audio understanding model as professional understanding models, and converting image, video and audio information into uniform text representation; s2, generating a structured text description through a cross-modal feature alignment mechanism, and forming a standardized text unit of multi-modal information; s3, based on a BERT generation embedded type, performing vectorization processing on a text query input by a user and a converted text; s4, performing similarity search in the vector database, and retrieving a text document most relevant to query; and S5, splicing the retrieved text and the user query, and inputting the spliced text and user query into the large language model to generate a final answer. According to the multi-modal retrieval enhancement generation method, a professional understanding model is innovatively introduced to serve as a different-modal encoder, the multi-modal information processing capacity of a large model is improved, and real cross-modal understanding is achieved.
Owner:GUANGZHOU BINGO SOFTWARE

Searching editing components based on text using a machine learning model

The present disclosure describes techniques for searching editing components based on text using a machine learning model. A plurality of visual embeddings indicative of a plurality of visual editing components is acquired by the machine learning model. The plurality of visual embeddings indicative of the plurality of visual editing components is projected into a common space by a first sub-model of the machine learning model. A text query input is received by a user. A text embedding indicative of the text query is generated. The text embedding is projected into the common space by a second sub-model of the machine learning model. At least one visual editing component among the plurality of visual editing components is determined based on the projected text embedding and the plurality of projected visual embeddings in the common space. Information indicative of the at least one visual editing component is displayed via a user interface.
Owner:LEMON INC(GB)

Landslide identification method and device based on SEEM-SAFPN model, medium and product

The invention discloses a landslide identification method and device based on an SEEM-SAFPN model, a medium and a product, and relates to the field of artificial intelligence. Firstly, a remote sensing landslide data set is obtained and preprocessed, and a multi-modal data set is constructed; each sample in the multi-modal data set comprises a remote sensing image, a text query, a non-text query and a corresponding mask; the SEEM model and the SAFPN module are combined to construct an SEEM-SAFPN model; the SEEM-SAFPN model comprises an encoder, a decoder and a prediction head; the encoder comprises a text encoder, an image encoder, an SAFPN module and a visual sampler; the multi-modal data set is divided into a training set, a verification set and a test set which are respectively used for training, verifying and testing the SEEM-SAFPN model, and the trained SEEM-SAFPN model is used as a landslide extraction model; the landslide extraction model is adopted to identify the landslide area in the remote sensing image, so that the landslide identification efficiency and accuracy can be remarkably improved.
Owner:SICHUAN INST OF LAND & SPACE ECOLOGICAL RESTORATION & GEOLOGICAL DISASTER PREVENTION

Cross-modal representation learning and retrieval method and system for grain production

The invention discloses a grain production-oriented cross-modal representation learning and retrieval method and system, and relates to the field of agricultural informationization, and the method comprises the steps: carrying out the multi-granularity semantic alignment of an image text pair in a grain production process based on an image-text bidirectional guidance fusion network, and obtaining a semantic segmentation image; performing image space decoupling and time sequence enhancement on a video text pair in a grain production process based on global semantic guidance to obtain a structured semantic image feature; constructing a text feature library and an image feature library; and determining a transmission plan matrix according to the modality of the to-be-retrieved data, generating query features of the to-be-retrieved data based on the transmission plan matrix, and outputting a text query result or an image query result by adopting a similarity measurement method according to the query features of the to-be-retrieved data, the text feature library and the image feature library. According to the method, deep fusion of cross-modal features can be realized, the semantic matching accuracy of the image and the text is improved, and rapid and accurate matching and retrieval between the image and the text are realized.
Owner:AGRI INFORMATION INST OF CHINESE ACAD OF AGRI SCI

Dialog agents with two-sided modeling

A central learning model is deployed as a user model and as an assistant model. Sensitive information utterances from a corpus of previously stored conversation language corresponding to user queries and chat agent responses thereto are used to train the user model to become an updated user model and to train the assistant model to become an updated assistant model, respectively. The user model provides user contexts corresponding to user queries to the assistant model and the assistant model provides assistant contexts corresponding to chat agent responses to the user model. During training, the user model does not provide plain-text queries to the assistant model and the assistant model does not provide plain-text responses to the user model. The updated assistant model may facilitate a federated training process produce an updated central model. An updated central model may be used to provide real-time chat agent responses to live user queries.
Owner:THE HONG KONG UNIV OF SCI & TECH

Multi-interface collaborative natural language data query method and system

The invention provides a multi-interface collaborative natural language data query method and system, and relates to the technical field of natural language processing.The method comprises the steps that text query data transmitted by a text input interface, voice query data transmitted by a voice input interface and touch instruction data transmitted by a touch selection interface are obtained; performing optimization processing on the voice query data and the touch instruction data to obtain a multi-interface fusion text set; analyzing the annotation data set based on preset natural language semantics to obtain a semantic analysis result; generating a unified query instruction through a digital information transmission technology; determining a data source corresponding to each piece of query field information in the semantic analysis result; according to the method, the initial query result is obtained according to the preset retrieval rule of each data source, so that accurate analysis and cross-interface unified execution of the natural language query intention are realized, and the accuracy, robustness and interactive experience of data query in a complex scene are improved.
Owner:BEIJING ALL VIEW CLOUD DATA TECH CO LTD

Retrieving digital images in response to search queries for search-driven image editing

Systems, methods, and non-transitory computer-readable media implements related image search and image modification processes using various search engines and a consolidated graphical user interface. For instance, one or more embodiments involve receiving an input digital image and search input and further modify the input digital image using the image search results retrieved in response to the search input. In some cases, the search input includes a multi-modal search input having multiple queries (e.g., an image query and a text query), and one or more embodiments involve retrieving the image search results utilizing a weighted combination of the queries. Some implementations involve generating an input embedding for the search input (e.g., the multi-modal search input) and retrieving the image search results using the input embedding.
Owner:ADOBE INC

Training method of video time positioning model, video time positioning method, equipment and medium

The invention relates to the technical field of computer vision, particularly provides a training method of a video time positioning model, a video time positioning method, equipment and a medium, and aims to solve the problem of large video time positioning error. In order to achieve the purpose, the model training method comprises the steps that multiple frames of images are sampled from a training video to serve as training data, the training data and a first preset query text are coded to obtain visual features and text query features, and a video time positioning model is trained based on the visual features and the text query features, obtaining a plurality of candidate answers, respectively calculating the relative advantage value of each candidate answer based on the real answer of the first preset query text and the plurality of candidate answers, adjusting the parameters of the video time positioning model based on each relative advantage value, and continuing to execute the step of sampling multiple frames of images from the training video. Therefore, the performance and accuracy of video time positioning can be improved.
Owner:PEKING UNIV +1

System for contextual searching using text search terms

A system for contextual matching video content based on multimodal metadata extraction generated by processing one or more scenes to extract metadata corresponding to multiple extraction modes, and an embedding model for each extraction mode wherein an aggregated embedding model responsive to said metadata embeddings for each mode formulates an aggregated embedding with an embedding extractor responsive to a text input with an embedding model coordinated with said embedding model wherein said embeddings are in the form of a vector, and a vector comparison processor for determining the distance between the query vector and a vector representing the aggregated embedding. The coordination between embedding models is established by training. The embedding extractor may accept a free-form text query and present one or more subqueries for embedding. A textual inversion engine may be provided to generate an image from the embeddings to provide feedback to a user. In this way a user can confirm the effectiveness of the text query. A text editor may be provided for a user to enter and to edit a query.
Owner:ANOKI INC

Machine-learned models for processing structural representations of therapeutic materials

PCT designated stageWO2025235348A1BiostatisticsCharacter and pattern recognitionStructural representationSequence processing
An example computer-implemented method for training a machine-learned sequence processing model to process queries over material structures includes obtaining a plurality of training inputs comprising a plurality of input sequences respectively corresponding to a plurality of representational schemas for representing material structures of a plurality of different materials, wherein each respective input sequence of the plurality of input sequences comprises a respective textual query and a respective sequence-based structural representation of a respective material, wherein each different respective sequence-based structural representation is characterized by a different respective representational schema; generating, using a machine-learned sequence processing model, a plurality of training outputs that respectively correspond to the plurality of training inputs; evaluating the plurality of training outputs; and updating, based on the evaluation of the plurality of training outputs, one or more parameters of the machine-learned sequence processing model.
Owner:GOOGLE LLC

LLM-based reasoning calculation method and device

The invention discloses an LLM-based reasoning calculation method, which comprises the following steps of: obtaining a context text queried from a knowledge base accessed by an LLM, and an intermediate reasoning result generated by reasoning calculation of the LLM aiming at a text unit contained in the context text; constructing a prefix tree based on the obtained context text; wherein each branch on the prefix tree respectively stores different context texts; calculating a cache income index corresponding to the text unit represented by the node on each branch of the prefix tree; wherein the cache revenue index is used for representing performance revenue which can be brought to reasoning calculation by multiplexing a cached intermediate reasoning result corresponding to the text unit; and sorting nodes on each branch of the prefix tree according to a sequence of values of the calculated cache income indexes from high to low, and after the nodes on each branch of the prefix tree are sorted, caching the intermediate reasoning result of the prefix tree in a cache space corresponding to the LLM.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Recursive deep text query method, system and equipment and storage medium

The invention discloses a recursive deep text query method, system and device and a storage medium, and the method comprises the steps: firstly creating an initial query and initializing a hierarchical depth, executing a first-layer query, carrying out evolution processing, and then carrying out retrieval; if the information density is insufficient, sub-query is generated to execute second-layer query, and retrieval is performed after second evolution processing; if the second-layer retrieval is contradictory, third-layer query is executed, retrieval is performed after timeline query and information source tracing verification are passed, and a third-layer retrieval result is finally output; through multi-level query and targeted evolution processing, query can be optimized step by step, more information dimensions can be covered, information gaps can be filled up, and retrieval comprehensiveness and accuracy can be improved; and a third-layer verification mechanism can effectively check contradictions by means of timelines and source tracing, so that the accuracy and reliability of results are guaranteed, and better retrieval results are provided for users.
Owner:ZHILU CLOUD (SHENZHEN) ARTIFICIAL INTELLIGENCE CO LTD

Multi-modal data processing and retrieval method

The invention particularly relates to a multi-modal data processing and retrieving method. The multi-modal data processing and retrieval method comprises the following steps: respectively carrying out depth feature extraction on image data and text data to generate an image vector and a text vector; carrying out interaction on the image vector and the text vector, and capturing semantic association information among multiple modes; mapping the vectors after interaction to a unified semantic space to realize semantic alignment among different modes; dividing the unified semantic space vector into fragments, storing the fragments in distributed nodes, and establishing a vector index; and generating a query vector by using an image or text queried by a user, carrying out parallel calculation on the similarity with a storage vector, and returning a retrieval result according to the similarity. According to the multi-modal data processing and retrieval method, the problems that the semantic difference between different modal data such as images and texts is large, the retrieval efficiency is low and storage is difficult to expand are solved, the accuracy and efficiency of multi-modal data retrieval are remarkably improved, and the method has remarkable technical advantages and wide application scenes.
Owner:INSPUR QILU SOFTWARE IND

API document recommendation method based on large language model

The invention discloses an API document recommendation method based on a large language model. The method comprises the steps of obtaining a plurality of text queries and a plurality of API documents, and presetting correlation values, so as to construct a reordering data set; training the constructed reordering model according to the reordering data set to obtain a trained reordering model; obtaining a to-be-tested text query, preprocessing all the API documents, and obtaining an initial sorting API document recommendation sequence of the to-be-tested text query according to the to-be-tested text query and the preprocessed API documents; and processing the initially sorted API document recommendation sequence and the to-be-tested text query by adopting the trained resorting model to obtain a resorted API document recommendation sequence of the to-be-tested text query. According to the method, the API document related to the development task is automatically recommended by understanding the query intention of developers, so that the consulting time of the developers is effectively shortened, and the development efficiency is improved.
Owner:ZHEJIANG UNIV