Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

169results about "Multimedia data querying" patented technology

Retrieval enhancement method based on multi-modal data fusion and modal perception

The invention relates to the technical field of information retrieval and generation, in particular to a retrieval enhancement method based on multi-modal data fusion and modal perception. According to the method, firstly, a dual-channel architecture is adopted to perform feature extraction and coding on a text and an image respectively, and mutually independent embedded representation spaces are constructed, so that high-quality collaboration and matching of cross-modal representation are realized; and a pseudo-pairing generation mechanism is introduced to effectively mine and reconstruct the existing non-paired data in the knowledge base. And designing a query modal perception and dynamic weighting mechanism for accurately controlling the fusion proportion of the image-text bimodal information in the retrieval stage so as to match the modal demand difference of different query contents. And further executing aggregation retrieval and reordering of the cross-modal information by using dynamic weighted fusion retrieval to generate a candidate set of multi-modal responses. According to the method, accurate matching and dynamic weight adjustment of the image-text content are realized, and the accuracy and expression integrity of the generated content are improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Cross-modal document information extraction method based on space-semantic alignment

The invention relates to a cross-modal document information extraction method based on space-semantic alignment, and belongs to the field of artificial intelligence, computer vision and natural language processing. According to the method, the spatial feature and semantic information bidirectional alignment model is designed, by constructing the spatial feature and semantic feature bidirectional alignment model, the document layout information can dynamically adjust attention distribution of text semantic features, meanwhile, semantic information reversely optimizes the spatial features, collaborative modeling of spatial layout and semantic information is achieved, and the document layout efficiency is improved. Therefore, the accuracy and robustness of complex document information extraction are improved. According to the method, a hierarchical cross-modal information extraction model is designed, through the hierarchical cross-modal information extraction model, the overall structure of a document is recognized on the global level, local key content is focused on the regional level, fine modeling is conducted on fine-grained texts and visual elements on the entity level, and accurate recognition of a cross-modal entity and the semantic relation of the cross-modal entity is achieved; and the generalization ability and applicability of information extraction are enhanced.
Owner:BEIJING INST OF COMP TECH & APPL

Nuclear power safety report examination method and system, electronic equipment and storage medium

The invention relates to the technical field of computers, and discloses a nuclear power safety report examination method and system, electronic equipment and a storage medium, and the method comprises the steps: receiving an examination request, carrying out the intention recognition and entity extraction based on a nuclear power field large language model, and disassembling an examination task into a structured subtask sequence based on an intention type and an entity attribute; performing multi-modal retrieval based on the structured sub-task sequence, and outputting a multi-source retrieval result; calling time-space association information of the evolutionary knowledge graph, and outputting knowledge support data in combination with the time sequence model; fusing a multi-source retrieval result and the knowledge support data, performing risk value calculation and uncertainty quantification through a deep learning model, and outputting a risk assessment result; and based on the risk assessment result, carrying out federal learning model parameter adjustment, carrying out uplink evidence storage on the review conclusion, the risk assessment result and the evidence chain, and outputting the review result. The method has the advantages of being efficient, accurate, dynamic, safe and closed-loop during examination.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Long-range visual question and answer and multi-modal reasoning task agent construction method

The invention provides a long-range visual question and answer and multi-modal reasoning task agent construction method. For input images and natural language questions, generating answers through orderly calling of an external tool set; the construction process of the intelligent agent is divided into three stages: a global planning stage: analyzing user requirements through a global planning navigator, and screening tool subsets; in the tool autonomous execution stage, visual details are supplemented, background knowledge is retrieved and accurate operation is carried out in the multi-step process by surrounding three types of tools through a tool autonomous executor, and a returned tool result is used for assisting the next step of decision making; in the response synthesis stage, the reasoning process is combed through a model response synthesizer, and user-friendly and clear-format responses are sorted; and finally, the model response synthesis module extracts and obtains answers from the reasoning track and outputs the answers. The problems of global planning insufficiency and visual context forgetting generally existing in complex task reasoning of an existing multi-mode large language model are solved.
Owner:RENMIN UNIVERSITY OF CHINA

Network model training method, data processing method, and apparatus

The present disclosure provides a network model training method, a data processing method, and an apparatus. The network model training method comprises: acquiring target sample data, wherein the target sample data comprises text sample data and image sample data; inputting the target sample data into a network model to be trained to obtain a sample recognition result; and adjusting a parameter of a text encoder on the basis of a text recognition result and first supervision data corresponding to the text recognition result, adjusting a parameter of an image encoder on the basis of an image recognition result and second supervision data corresponding to the image recognition result, and a hybrid image-text recognition result and third supervision data corresponding to the hybrid image-text recognition result, and adjusting a parameter of a hybrid encoder on the basis of the hybrid image-text recognition result and the third supervision data corresponding to the hybrid image-text recognition result to obtain the trained network model formed by the text encoder, the image encoder, and the hybrid encoder.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Multi-dimensional knowledge base construction and query method based on multi-modal large model

The invention relates to a multi-dimensional knowledge base construction and query method based on a multi-modal large model, and belongs to the technical field of AI large model application. The method comprises the following steps of: analyzing a text file containing text, table and picture contents; selecting a pre-trained large language model and a multi-modal large model, generating text abstracts and picture abstracts by utilizing the models, and enabling storage paths of the picture abstracts to correspond to storage paths of pictures one by one; selecting a text embedding model to construct a vector database and a retriever; and constructing a retrieval chain based on the multi-modal large model. According to the method, construction from a single text knowledge base to a multi-dimensional knowledge base is achieved, a multi-dimensional information retrieval method is provided, the problem of construction of the multi-dimensional knowledge base based on the vectorization technology is solved to a certain extent, and the large model output quality based on the retrieval enhancement generation technology is remarkably improved.
Owner:HONGHE POWER SUPPLY BUREAU OF YUNNAN POWER GRID

Multi-modal data retrieval, generation and synthesis method and system based on artificial intelligence driving

The invention relates to the technical field of artificial intelligence, in particular to a multi-modal data retrieval, generation and synthesis method and system based on artificial intelligence driving, and the method comprises the steps of multi-modal feature extraction, cross-modal alignment, feature fusion, multi-modal retrieval and sorting and multi-modal generation. Different deep learning models are adopted to carry out feature extraction on multi-modal data and convert the multi-modal data into vectors, through cross-modal alignment, different modal feature vectors are mapped to the same vector space, through feature fusion, correlation weights are calculated through an attention mechanism, weighted summation is carried out on fusion features, and through multi-modal retrieval and sorting, multi-modal data are obtained. According to the method, the similarity between a query vector and candidate data is calculated and sorted, and finally, through multi-modal generation, retrieval knowledge is used as external knowledge to be fused into a generation model, and a multi-modal synthesis answer is generated, so that information in multi-modal data can be effectively integrated and utilized.
Owner:BEIJING SGITG ACCENTURE INFORMATION TECH CO LTD +1

Cross-modal conference information association retrieval method and system and medium

The invention discloses a cross-modal conference information association retrieval method and system and a medium, and relates to the technical field of artificial intelligence, and the method comprises the following steps: carrying out feature extraction on obtained multi-source heterogeneous data to obtain multi-modal features, and uniformly mapping the multi-modal features to a first feature space of a preset dimension; in the first feature space, cross-modal deep fusion processing is performed on the multi-modal features, and a joint embedding space with consistent semantics is constructed according to the cross-modal deep fusion processing; constructing a vector index database based on a multi-modal feature vector in the joint embedding space, receiving a natural language query and mapping the natural language query to the joint embedding space, executing two-stage retrieval, and then obtaining a semantic fusion score based on calculated semantic fusion scores; multiplying a time sequence reward value based on the query time deviation and a dynamic reward value based on the core word matching degree to obtain a dynamic fusion score, and performing fusion sorting on the candidate set to obtain a final sorting result and an associated retrieval result; according to the method, refined sorting of the retrieval results is realized.
Owner:UNIV OF SCI & TECH OF CHINA

RAG-based pdf intelligent retrieval and generation method and system

The application discloses a kind of PDF intelligent retrieval and generation method and system based on RAG, by obtaining the document data of input, using the classification model established in advance to parse document data, extract text content and image content to form first data set;Using deep learning model to the image content in first data set carries out feature extraction, while the text content in first data set applies natural language processing technology to carry out semantic analysis, obtains multimodal feature set;According to multimodal feature set, application information integration algorithm is uniformly encoded and is handled to generate second data set, if detecting the integrity of fusion feature vector in second data set is lower than preset threshold value, then supplementary context semantic analysis fills in missing information;Using preset index construction mechanism to the clustering processing of fusion feature vector in second data set, generates the retrieval index library containing classification index structure.The application improves the accuracy and comprehensiveness of document retrieval.
Owner:HUNAN ZHIXUE YOUKE INFORMATION TECHNOLOGY CO LTD +1

Visual compression and retrieval method and device of document, equipment and storage medium

The invention discloses a document visual compression and retrieval method and device, equipment and a storage medium, and relates to the technical field of computers, the method comprises the following steps: obtaining a to-be-processed document page image, segmenting the image into a plurality of image blocks, and determining the structure category and the structure importance score of each image block; obtaining a plurality of structure regions based on structure category and spatial position aggregation, and distributing a preset number of compression tokens for each region in combination with structure category weights and importance scores; generating structure anchor point tokens corresponding to the regions by the compressed tokens to form a set; receiving a query request, converting the query request into a query vector, and performing retrieval in the anchor point token set to obtain a target structure region; and performing local decoding reconstruction based on the compressed token of the target region, and outputting a region image or a structure mask. According to the method, through structure-guided self-adaptive compression and fine-grained retrieval, the long document processing efficiency is greatly improved, and the compression effect and the retrieval accuracy are both considered.
Owner:BEIJING DIGITAL CHINA CLOUD COMPUTING CO LTD

Water conservancy project file retrieval method and system

The invention relates to the technical field of electronic data processing, and discloses a hydraulic engineering archive retrieval method and system.The hydraulic engineering archive retrieval method comprises the steps that a hydraulic engineering archive containing text data and image data is obtained, and the text data and the image data are mapped into an initial text feature vector and an initial image feature vector; the initial text feature vector and the initial image feature vector are projected to a shared semantic subspace, cross-modal representation is obtained, a joint loss function comprises a local feature alignment loss item and a semantic structure maintenance loss item, for any query request, the query request is converted into a query vector, and in the shared semantic subspace, the local feature alignment loss item and the semantic structure maintenance loss item are obtained. And sorting by calculating the cosine similarity of the query vector and the cross-modal representation of all archives in the archive library, and outputting a sorting result as retrieval. According to the method, fine-grained semantic alignment between the local key information of the image and the core terms of the text is realized, the understanding depth and accuracy of the core associated content of the image and text are improved, and the file retrieval precision is improved.
Owner:LUOYANG WATER ECOLOGY INVESTMENT & DEVELOPMENT CO LTD

Image-text retrieval method and system based on multi-modal fusion and depth spectral clustering

The invention discloses an image-text retrieval method and system based on multi-modal fusion and depth spectral clustering, and relates to the technical field of data retrieval, and the method comprises the steps: extracting image training features and text training features; splicing to obtain a global feature; constructing local image training features and local text training features; performing alignment processing on the local image training features and the local text training features; fusing the global features and the aligned local image training features and local text training features; performing clustering analysis on the normalized features after fusion feature normalization to obtain a plurality of clustering centers; and extracting query features of the query sample, and determining a retrieval result according to the cosine similarity. According to the method, the characterization capability and robustness of the model to the multi-modal data are improved through data-level fusion and local structure constraint, self-supervised alignment of the image and text modal data is realized through the positive sample pair and the negative sample pair, the clustering center of the image-text data is learned through depth spectral clustering, and new sample data are quickly adapted.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Index recovery method and device for streaming media data and storage medium

The invention discloses an index recovery method and device for streaming media data and a storage medium, and the method comprises the steps: persistently storing a data frame to a data storage area, caching an index to an index cache area, and sending the index to a slave storage device for backup storage; if the preset index flashing condition is met, persistently storing the cached index into an index storage area; after the main storage device is abnormal, the main storage device returns to normal, a backup index in the slave storage device is obtained, and the backup index is persistently stored in an index storage area; and if the index corresponding to the streaming media in the index storage area is still not completely recovered, reading the index of the data frame in the data storage area, and persistently storing the read index in the index storage area. The index recovery is carried out according to the backup index in the slave storage device, the index recovery speed can be increased, the information of the recovered index is more complete, then the index of the data frame in the data storage area is read for index recovery, and the index integrity of each data frame is ensured.
Owner:ZHEJIANG DAHUA TECH CO LTD

Query method and electronic device

Embodiments of the present application provide a query method and an electronic device. The query method comprises: in response to querying a document by using a first query statement, obtaining a query word library and a document library; the query word library comprises an original query statement and a rewritten query statement; the document library comprises original documents and rewritten documents associated with each query statement; matching the first query statement with the query word library; in the case that the query word library contains the first query statement, obtaining a target document associated with the first query statement from the original documents and / or the rewritten documents; in the case that the query word library does not contain the first query statement, rewriting the first query statement by using a pre-trained online query statement rewriting model to obtain a second query statement; and obtaining a target document associated with the second query statement from the original documents and / or the rewritten documents. The present application can enhance the processing capability of the query system for complex query statements, significantly expand the document recall dimension, and optimize the query effect.
Owner:RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD

Multi-modal information retrieval and analysis method for internet real-time information flow

The invention provides a multi-modal information retrieval and analysis method for internet real-time information streams, which comprises the following steps of: acquiring multi-modal data in the internet, including texts, images, videos and audios, and performing timestamp calibration on each modal data to form a multi-modal data set; based on the multi-modal data set, respectively extracting semantic features of each modal through a specific feature extractor, including semantic vectors of texts, visual features of images, spatial-temporal features of videos and audio features of audios; according to the multi-modal information retrieval and analysis method and system for the internet real-time information flow, multi-modal feature alignment is achieved by building a cross-modal joint embedding space, the retrieval precision is improved by adopting a three-stage retrieval assembly line, the calculation performance is optimized by combining model quantification and a heterogeneous calculation platform, and the multi-modal information retrieval and analysis method and system for the internet real-time information flow are achieved. The method has the advantages that the cross-modal retrieval accuracy is improved, the real-time processing capability is enhanced, and the calculation efficiency is optimized.
Owner:北京圆璟科技有限公司

Multimedia resource loading method and related device

The present disclosure provides a multimedia resource loading method, comprising: in response to a trigger operation of requesting to display a target page, obtaining a resource mapping table corresponding to the target page; determining an identifier of a target multimedia resource to be loaded on the target page; determining a storage path of the target multimedia resource according to the resource mapping table corresponding to the target page based on the identifier of the target multimedia resource; and obtaining and loading the target multimedia resource based on the storage path of the target multimedia resource. Based on the above multimedia resource loading method, the present disclosure further provides a multimedia resource loading device, an electronic device, a storage medium and a program product.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Hydroelectric equipment knowledge graph retrieval enhancement method and device based on multi-modal embedding

The invention provides a hydroelectric equipment knowledge graph retrieval enhancement method and device based on multi-modal embedding, and belongs to the technical field of intelligent perception and multi-modal semantic understanding of water conservancy and hydropower engineering. The method comprises the following steps: on the basis of related documents and pictures of hydroelectric equipment, constructing a multi-modal knowledge graph by utilizing a large language model; performing vectorization representation on pictures and text information in the multi-modal knowledge graph, and mapping the pictures and the text information to a unified semantic space to generate a vector database used for retrieving the multi-modal knowledge graph; and the vector database is utilized to retrieve the multi-modal knowledge graph. According to the method, deep fusion and semantic alignment of multi-source heterogeneous data such as images and texts can be realized, the method can be widely applied to fields such as hydropower station electromechanical equipment maintenance, hydraulic structure safety inspection and scheduling operation and maintenance knowledge management, and the equipment identification precision, the retrieval response speed and the knowledge service intelligence level can be remarkably improved.
Owner:GUODIAN DADUHE HOUZIYAN HYDROPOWER CONSTR CO LTD

Apparatus and method for providing telecommunication routing data

Systems and methods are disclosed to provide routing data associated with a destination phone number to a service provider. The service provider sends a first query to a first computing node of a data provider for routing data associated with a phone number. The first computing node provides the routing data if it is available from the data provider. Otherwise, the first computing node identifies a second data provider that is capable of providing the routing data. The first computing node then obtains the routing data from the second data provider and returns it to the service provider.
Owner:NETNUMBER COM

Apparatus and method for providing telecommunication routing data

Systems and methods are disclosed to provide routing data associated with a destination phone number to a service provider. The service provider sends a first query to a first computing node of a data provider for routing data associated with a phone number. The first computing node provides the routing data if it is available from the data provider. Otherwise, the first computing node identifies a second data provider that is capable of providing the routing data. The first computing node then obtains the routing data from the second data provider and returns it to the service provider.
Owner:NETNUMBER COM

Data retrieval method and apparatus, data processing method and apparatus, device, and medium

Provided are a data retrieval method and device, a data processing method and device, equipment and a medium, which are related to the technical field of computers and can be used for data retrieval and data clustering. The implementation method comprises the following steps: in response to receiving a to-be-retrieved vector, determining at least one to-be-retrieved storage medium corresponding to the to-be-retrieved vector; for each to-be-retrieved storage medium in the at least one to-be-retrieved storage medium, using a storage controller corresponding to the to-be-retrieved storage medium to extract a first number of sample vectors from at least one sample vector stored in the to-be-retrieved storage medium, the similarity of the first number of sample vectors to the to-be-retrieved vector being higher than the similarity of other sample vectors in the at least one sample vector to the to-be-retrieved vector; and determining a retrieval result corresponding to the to-be-retrieved vector based on the first number of sample vectors from each of the at least one to-be-retrieved storage medium.
Owner:VASTAI TECH (SHANGHAI) INC

Context information management method and device, equipment, medium and program product

The invention discloses a context information management method and device, equipment, a medium and a program product, and relates to the technical field of data processing, and the method comprises the steps that a first component command is acquired, and the first component command carries a collection range of context information; the execution of the context component is triggered based on the first component command to obtain a first vector library, the first vector library is in one-to-one correspondence with the context type of the context information, and the context component is configured to be provided with a context acquisition module, a context processing module, a context storage module and a context retrieval module; the context storage module is used for storing the processed context information based on the context type of the processed context information, the context retrieval module is used for providing a retrieval tool aiming at the first vector library, and the retrieval tool is configured to be provided for the context processing module to call or provide a first object call of a first component command. According to the method, the context information in the first vector library can be efficiently and accurately obtained.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Multimedia file retrieval processing method and device

The invention discloses a retrieval processing method and device for multimedia files, and belongs to the technical field of communication. The method is applied to the electronic equipment and comprises the following steps: acquiring keyword information of multimedia file retrieval; searching target content information corresponding to the keyword information in a multimedia library; wherein the multimedia library records a corresponding relationship between a plurality of extraction fields of each multimedia file and content information of the extraction fields in the multimedia file; and displaying the target content information.
Owner:VIVO MOBILE COMM CO LTD

A material rendering method and apparatus

The embodiment of the present application provides a kind of material rendering method and device, it is related to data processing technical field, find second material;If second material is not found, remove first material from first material list, and render first material;If second material is found, in the case where the content of first material and the content of second material are different, update first material to second material, render first material, and render second material, in the case where first order and second order are different, third material is inserted to first order in first material list, render third material;Material at next order is used as new first material;Material in second material list after current first order is added to last material in first material list, and the material added to first material list is rendered.The scheme provided by the embodiment of the present application can reduce the consumption of processing resources in material rendering process.
Owner:HANGZHOU EZVIZ SOFTWARE CO LTD

System and method for managing an enhanced performance network using role-based agents

Aspects of the subject disclosure may include, for example, a device, including: a processing system including a processor; and a memory that stores executable instructions that, when executed by the processing system, facilitate performance of operations of: training a role-based agent to handle queries from users of a network; training a role-based foundation model to retrieve data relevant to a role of a user; receiving a query from the user; and providing the query to the role-based agent, wherein the role-based agent uses an associated role-based foundation model to process the query, collect relevant data from the network, and formulate an answer to the query. Other embodiments are disclosed.
Owner:AT&T INTELLECTUAL PROPERTY I L P +1

Method, apparatus, device and storage medium and program product for information processing

The embodiment of the disclosure provides a method, apparatus, device, storage medium and a program product for information processing. The method comprises: obtaining, in response to receiving a user query from a user to a digital assistant from a user, auxiliary visual information acquired by a visual acquisition tool associated with the user based on the user query, an acquisition time of the auxiliary visual information being adjacent to an initiation time of the user query; and determining a reply corresponding to the user query based on the acquired auxiliary visual information and the user query with a machine learning model.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Multi-source education data searching method, device and equipment and storage medium

The application discloses a multi-source education data search method and device, equipment and a storage medium, and the method comprises the steps of: obtaining education information from a plurality of heterogeneous data sources through a hybrid search strategy, and performing recursive data association on the education information to generate an extended information set; performing data screening on the extended information set based on a pre-trained model to obtain a screened extended information set; receiving a user query request, using a multi-modal hybrid retrieval model to search information in the screened extended information set, and obtaining a search result. Compared with the prior art, the application eliminates information silos and improves the education information search matching accuracy by obtaining education information from a plurality of heterogeneous data sources through a hybrid search strategy, performing recursive data association on the education information to generate an extended information set, receiving a user query request, and using a multi-modal hybrid retrieval model to search information in the screened extended information set.
Owner:SHENZHEN ZHISI EDUCATION TECH CO LTD

Multi-modal document slicing method and related device

The invention discloses a method for slicing a multi-modal document, which is applied to the technical field of artificial intelligence (AI). The multi-modal document slicing method comprises the following steps: segmenting text data in a multi-modal document to obtain a plurality of text segments; moreover, data (such as images or tables) of other modalities in the multi-modal document are processed through the model, and contents included in the data of other modalities can be described in a text form. Therefore, correlation between the text fragments and other modal data can be realized by calculating the correlation degree between the plurality of text fragments obtained by segmentation and the description texts of other modal data, so that on the basis of slicing different modal data in the document, an association relationship is established for the sliced data, and the data processing efficiency is improved. And the fine-grained retrieval requirement of the user for the multi-modal document can be met.
Owner:HUAWEI TECH CO LTD

Information retrieval method and device, electronic equipment, storage medium and program product

Some cases involve retrieval technology field, information retrieval method, device, electronic equipment, storage medium and program product are disclosed, the method comprises: obtaining first information;Obtain the first feature vector of the first information and the first discrete code word of the first information;Contrast the first discrete code word and the second discrete code word, obtain the first contrast result, the second discrete code word corresponds to the second information;Filter the second information by using the first contrast result, obtain the third information;Contrast the first feature vector and the second feature vector, obtain the second contrast result, the second feature vector corresponds to the third information;Filter the third information by using the second contrast result, obtain the fourth information, the fourth information is used to represent the retrieval result of the first information. Through the implementation of the information retrieval method, the retrieval accuracy and the overall retrieval efficiency can be improved.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Video conference transcript querying using artificial intelligence

Techniques for video conference transcript querying using artificial intelligence are provided. In an example method, a video conference provider joins a client device to a video conference. The video conference provider receives a deletion election. The video conference provider receives an audio stream from the client device and generates a portion of a transcript of the video conference. Prior to the video conference concluding, the video conference provider processes the portion of the transcript to configure a large language model (“LLM”) to respond to queries based on the video conference. The video conference provider receives a query relating to the video conference and causes the LLM to process the query and the portion of the transcript. The video conference provider outputs a response, generated by the LLM, to the query. The video conference provider then deletes the portion of the transcript based on the deletion election.
Owner:ZOOM COMMUNICATIONS INC