Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

278 results about "Inverted index" patented technology

In computer science, an inverted index (also referred to as a postings file or inverted file) is a database index storing a mapping from content, such as words or numbers, to its locations in a table, or in a document or a set of documents (named in contrast to a forward index, which maps from documents to content). The purpose of an inverted index is to allow fast full-text searches, at a cost of increased processing when a document is added to the database. The inverted file may be the database file itself, rather than its index. It is the most popular data structure used in document retrieval systems, used on a large scale for example in search engines. Additionally, several significant general-purpose mainframe-based database management systems have used inverted list architectures, including ADABAS, DATACOM/DB, and Model 204.

Data reordering retrieval method and system based on RAG

The invention provides a data reordering retrieval method and system based on RAG, and the method comprises the steps: carrying out the dynamic semantic partitioning processing of an original document, generating corresponding document blocks, storing the document blocks in a vector database, constructing a hierarchical index, analyzing a received query request, extracting keywords in the query request, and carrying out the retrieval of the query request. Boolean keyword matching is carried out through an inverted index in the hierarchical index, semantic retrieval of keywords is carried out in a vector database, an initial candidate set is generated after multi-source retrieval results are fused, a query request and the initial candidate set are input into a generative reordering model, and a query result is obtained; according to the method, the initial candidate set is selected, the corresponding score is generated through the cross-modal attention mechanism in the generative reordering model, the initial candidate set is reordered according to the score, the final retrieval result is obtained, and the reordered document blocks are displayed, so that the overall retrieval speed is increased, the limitation of a single retrieval mode is avoided, and the accuracy of the retrieval result is improved.
Owner:TAIJI COMPUTER CORPORATION LIMITED

Permanent incremental backup method based on S3 bucket-level virtual snapshot

The invention discloses a permanent incremental backup method based on an S3 bucket-level virtual snapshot, and relates to the technical field of data storage and backup, the method comprises the following steps: initializing a backup system and starting a backup agent program, an index management service and an S3 object storage service; complete backup is executed, and a VID and S3 bucket-level virtual snapshoot are generated; efficiently managing a mapping relation between a file path and a snapshot identifier through a B + tree and an inverted index by utilizing a MongoDB index service; incremental backup is executed, and file changes are accurately recognized through Hash comparison and modification time comparison; a storage position is optimized by adopting an incremental storage and deduplication storage mechanism and an IncDSI optimization algorithm; executing backup chain integration and cleaning; and performing data recovery based on the VID of the file and the directory. According to the method, the backup efficiency is remarkably improved, the storage overhead is reduced, the data recovery process is optimized, and the method is suitable for data backup and management of a large-scale file system and has important application value and prospect.
Owner:NANJING UNARY INFORMATION TECH

Intelligent search engine system and method based on NLP and vector hybrid retrieval

The invention relates to the technical field of natural language processing, in particular to an intelligent search engine system and method based on NLP and vector hybrid retrieval, and the system comprises a query analysis module which is used for receiving a query statement input by a user and obtaining structured query information based on the query statement; the symbiotic index module comprises a sparse extension index unit which is used for constructing a sparse inverted index table based on keywords in the structured query information; the vectorization hypergraph index unit is used for forming a hypergraph index based on the business data; the recall arrangement module is used for acquiring a candidate object set according to the structured query information; the constraint rearrangement module is used for calculating a joint priority score between query and candidate objects based on the candidate object set, and obtaining a sorting result under constraint conditions of meeting a supplier proportion, a category proportion and a price interval; and the evidence generation module is used for generating evidence information based on the sorting result and outputting the evidence information to the user interface.
Owner:BEIJING JINGNENG TENDERING & COLLECTIVE PROCUREMENT CENT CO LTD

Electronic archive intelligent retrieval method and system based on block chain

The invention discloses an electronic archive intelligent retrieval method and system based on a block chain, and relates to the technical field of archive retrieval, and the method comprises the steps: extracting text features of an electronic archive, and generating metadata; generating a content hash value from the electronic file original text; writing the metadata, the content hash value and the ABE access strategy into the block chain smart contract; constructing a reverse index based on the metadata semantic tag, wherein an index entry is associated with a block chain storage address; calculating a root hash and anchoring the root hash to the block chain; analyzing a keyword and a digital identity certificate in the user retrieval request; calling an intelligent contract to verify whether the user attribute accords with the ABE access strategy of the target file; retrieving an encrypted hash list matched with the file in the distributed index; acquiring the encrypted file fragments from the distributed storage system; verifying data integrity; and combining the fragments to generate a final retrieval result. The method has the advantages that through deep coupling of the block chain and attribute encryption, an electronic archive management system considering security and intelligent retrieval is constructed.
Owner:BEIJING RUIYUN ARCHIVES MANAGEMENT CO LTD

Method and system for retrieving DOCX document content based on keywords

The invention belongs to the technical field of text processing, and particularly relates to a method and system for retrieving DOCX document content based on keywords, which comprises the following steps: analyzing an Office Open XML structure of a DOCX document, combining with multi-dimensional features such as style names, and utilizing a title classification score model to accurately distinguish a title and a text, so that a semantic hierarchical structure of the document is effectively reserved; and secondly, a multi-level semantic extension mechanism is introduced, and a Sension-BERT, a HowNet knowledge base and a Word2Vec model are fused, so that intelligent extension of synonyms and synonyms of keywords is realized, and the recall rate and semantic understanding ability of retrieval are remarkably improved. And in addition, a BM25 model is combined with paragraph length normalization and structure position weight to calculate a correlation score, so that retrieval results are sorted more accurately and reasonably. The construction of the reverse index is combined with the position coding and compression optimization strategy, and the retrieval efficiency and the storage performance are both considered.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Retrieval method and device based on static word embedding, computer equipment and medium

The invention relates to a retrieval method and device based on static word embedding, computer equipment, a computer readable storage medium and a computer program product. The method comprises the steps of performing word segmentation on an original text to obtain a first word segmentation result, training a static word embedding model by utilizing the first word segmentation result to obtain word vectors, and generating a synonym word library; the method comprises the following steps: establishing a full-text inverted index by utilizing an original text, and expanding query words by utilizing a synonym library in a retrieval stage; encoding the original text into a semantic vector by using a semantic generation model, and constructing a vector index based on the semantic vector; based on a to-be-queried text in a user query request, performing retrieval by using the full-text inverted index to obtain a first candidate document, and performing retrieval by using the vector index to obtain a second candidate document; performing fusion processing on the first candidate document and the second candidate document to obtain a target candidate document; and inputting the target candidate document into the text generation model to obtain a retrieval result. By adopting the method, the accuracy of text retrieval can be improved.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Electronic archive intelligent retrieval method and system

The invention discloses an intelligent retrieval method and system for electronic archives, and belongs to the field of electric digital data processing.The intelligent retrieval method for the electronic archives comprises the following steps that search content input by a user is received, spelling errors are automatically corrected, and related words are expanded; preliminarily screening the candidate text electronic archive set based on a reverse index mechanism to obtain an initial sorting result, performing correlation resorting and extension recall on the initial sorting result through a semantic similarity model, and outputting a final sorting result; and after determining the target text electronic archive in the final sorting result, according to the uploading information of the target text electronic archive, the method has the beneficial effects that the non-text electronic archives are matched through the determined uploading information of the target text electronic archive, and the non-text electronic archives are sorted according to the association strength, so that the efficiency of sorting the non-text electronic archives is improved. The non-text format content can be identified during keyword retrieval; and the tendency of the user is predicted, so that the use experience of the user in downloading or checking the electronic file is improved.
Owner:SHANDONG YANGMING NETWORK INFORMATION TECHNOLOGY CO LTD

Data segmentation and query processing system under distributed architecture

The invention relates to the technical field of computer data processing, and discloses a data segmentation and query processing system under a distributed architecture. The system comprises a data segmentation module which divides data blocks based on a self-adaptive partitioning algorithm; the distributed storage module is used for distributing data blocks by using a consistent Hash algorithm and dynamically adjusting the distribution of virtual nodes; the index construction module is used for generating a global index table through a multi-layer inverted index algorithm; the query optimization module is used for converting query statements by utilizing a query rewriting algorithm; and the result aggregation module is used for processing the sub-query results by adopting a parallel aggregation algorithm. The system can dynamically process data according to data characteristics, optimizes the data storage and query process, effectively improves the data processing efficiency and query performance of a distributed system, and is suitable for scenes such as big data storage and analysis.
Owner:ZHEJIANG JINGJING TECH CO LTD

Standard address pipeline data association management method and system based on multistage collaborative coding

The invention discloses a standard address pipeline data association management method and system based on multi-level collaborative coding, belongs to the technical field of government affair digitization, and aims to solve the technical problem of how to generate and manage standard address data in government affair business, realize automatic triggering and standardized circulation of cross-department business data and improve the management efficiency. According to the technical scheme, the method comprises the steps of cross-department business triggering: automatically triggering a road and courtyard address code generation process through construction project land pre-examination and construction project planning permission business events; a standard road name is generated based on road pre-examination data, a road unified identification code is synthesized according to a 4 + 6 + 9 segmentation rule, a house number is generated according to planning permission data, and a courtyard unified identification code is generated in combination with a naming system; the house prediction data triggers the generation of a house unified identification code; and data dynamic association: constructing a reverse index based on the address coding feature vector, and realizing millisecond-level accurate matching of the address and population information through a distributed query engine.
Owner:浪潮智慧城市科技有限公司

Credible query configuration method and system for block chain PB-level multi-modal data

The invention belongs to the technical field of block chains, and discloses a trusted query configuration method and system for block chain PB-level multi-modal data, and the method comprises the steps: obtaining multi-modal uplink data, and extracting metadata of the multi-modal data to construct an in-block index; constructing a block index table based on the block ID, logically organizing the block index table, and adaptively and dynamically adjusting the block index table; constructing an index partition table used for storing block index table identifiers and addresses; constructing a coarse-grained index model, quickly positioning a block index table in an index partition table according to query conditions, searching a reverse index in the block index table, obtaining a block ID (Identity) and a storage position, retrieving an index in a block, querying a specific data record, verifying a Hash path, and realizing credible query of data. According to the method, a decentralized self-adaptive multi-level index mechanism is adopted, decentralized query performance is optimized, complex query and traceability query operations such as connection and aggregation are supported, and effectiveness and credibility of query results are guaranteed through a credible query algorithm.
Owner:JINAN SHENGAN INFORMATION TECH CO LTD

Hotel room type matching method and device

The invention discloses a hotel room type matching method and device, and the method comprises the steps: obtaining the original information of a newly-added hotel and room type, carrying out the preprocessing of the original information, and generating a standardized hotel and room type feature data set; constructing a multi-level feature vector based on the standardized feature data set; screening and generating a candidate matching set in combination with an inverted index and a locality sensitive hashing algorithm; a matching score matrix is calculated through fusion of the rule matching model, the machine learning model and the deep learning model; generating a final matching conclusion through Bayesian uncertainty estimation and threshold dynamic adjustment; and optimizing the matching model based on the matching conclusion and the business feedback data. Through multi-dimensional feature extraction, multi-model fusion matching and confidence evaluation, the problem of accuracy of hotel room type cross-supplier matching is solved, and the matching efficiency and robustness are improved.
Owner:GUIZHOU YOUTEYUN TECH CO LTD

Text content index automatic identification method based on semantics

The invention discloses a semantic-based text content index automatic identification method, which relates to the technical field of information retrieval, and comprises the following steps: initializing sparse projection and LSH signature, performing iterative optimization by using a Lagrange duality form and gradient update, adjusting hash digits, obtaining a fragment index through a k-d tree, and constructing an inverted index. According to the method, compression is performed through Delta coding, an index map is constructed based on Jaccard similarity, compression is performed through WebGraph, CSNMF is used in combination with Z-Laplacian regularization, a low-rank basis matrix and a low-rank coding matrix are generated, a compressed inverted index is reconstructed after iterative optimization, and reconstructed inverted index entries are generated. According to the method, through multi-resolution hash table initialization, joint feature optimization, local adaptive quantization and low-rank index reconstruction, the semantic expression ability and the compression effect of an index structure are improved, the index precision and efficiency are improved, and intelligent identification of index content is achieved.
Owner:BEIJING GEPU TECHNOLOGY CO LTD

Alarm information processing method and device and storage medium

The invention discloses an alarm information processing method and device and a storage medium, and relates to the technical field of data processing.The method comprises the steps that an unstructured alarm text is converted into an accurate keyword set through word segmentation processing, the limitation of traditional simple text matching is broken through, alarm core semantic information is effectively extracted, and redundant interference is reduced; according to the method, the mapping relation between the keyword and the alarm identifier is established and stored in the database, so that an efficient retrieval structure based on the inverted index is realized, the overhead of full-text scanning is avoided, and the retrieval speed of massive alarm data is remarkably improved; according to the method, the target alarm text can be more accurately positioned based on semantic relevance of the keywords, and the rigidity of traditional fixed condition retrieval is overcome, so that the practicability of alarm information retrieval is synchronously optimized in the aspects of efficiency and precision. The technical problem of low retrieval efficiency in a traditional retrieval mode is solved, and the technical effect of improving the retrieval speed of mass alarm data is achieved.
Owner:ZHENGZHOU YUNHAI INFORMATION TECH CO LTD

Retrieval method and system suitable for different data densities

The invention provides a retrieval method and system suitable for different data densities, and the method comprises the steps: distributing each data point to a corresponding hash bucket based on a locality sensitive hash algorithm, and forming an initial inverted index based on an identifier of the hash bucket and the data points; processing data points in each hash bucket based on a hierarchical navigable small world insertion algorithm to form an HNSW graph, and forming an inverted index containing an HNSW graph index; obtaining a candidate bucket set of the query point; obtaining an HNSW graph corresponding to each candidate bucket in the candidate bucket set based on the constructed inverted index; retrieving from the HNSW graph corresponding to the candidate bucket based on a hierarchical navigable small world search algorithm to obtain preliminary retrieval results, and determining a global retrieval result according to the preliminary retrieval results and returning the global retrieval result. According to the method, efficient retrieval of dense vectors and sparse vectors can be carried out uniformly, the retrieval efficiency and precision are improved, the universality is high, and the complexity and development cost of the system are reduced.
Owner:THE THIRD RES INST OF MIN OF PUBLIC SECURITY

Intelligent Agent collaborative decision-making system for environment monitoring

The invention relates to the technical field of water quality monitoring and analysis, in particular to an intelligent Agent collaborative decision-making system oriented to environment monitoring. Comprising the steps that original data including water quality parameters and environment data are collected in real time; performing noise elimination, missing value filling and standardization processing on the collected original data, and outputting preprocessed data; the method comprises the following steps: integrating multi-source data including historical monitoring data, pollution source information and treatment cases, performing vectorization coding on the multi-source data by adopting a BERT model, and constructing a retrieval architecture based on hierarchical indexes and inverted indexes of a hierarchical navigable small world to obtain a dynamic water quality knowledge base; and on the basis of a retrieval enhancement generation technology, associating the preprocessed data with the dynamic water quality knowledge base, predicting a water quality change trend through a time sequence prediction model, and generating an abnormality diagnosis report and a closed-loop treatment scheme. According to the invention, analysis and early warning of water quality are realized through collaborative decision-making of the intelligent Agent.
Owner:ZHEJIANG GUANGCHUAN ENG CONSULTING CO LTD

Distributed aggregation retrieval method for law and regulation text database

The invention relates to the technical field of aggregation retrieval, in particular to a distributed aggregation retrieval method for a law and regulation text database, which comprises the following steps: collecting and integrating law and regulation text data, preprocessing the law and regulation text data, removing useless punctuations and blank characters, and generating a preprocessed text set. According to the method, data integration and preprocessing are carried out on the regulation text, useless punctuations and redundant blank characters are effectively removed, the purity of text data is guaranteed, and the accuracy of subsequent word segmentation and word frequency statistics is improved; the implementation of the word segmentation and word frequency statistical process is helpful for identifying key feature vocabularies in the regulatory text, so that the accuracy of the subsequent inverted index construction stage is enhanced; by preliminarily constructing the inverted index and implementing similar item merging and low-frequency vocabulary removing operation, the index scale and redundant interference are reduced, and the index query efficiency and accuracy are improved.
Owner:NANJING XIAOZHUANG UNIV

Intelligent dialogue knowledge base optimization method based on multi-path recall fusion and self-updating

The invention discloses an intelligent dialogue knowledge base optimization method based on multi-path recall fusion and self-updating, and relates to the technical field of artificial intelligence. The method comprises the following steps that A, a multi-path parallel recall engine is constructed, a knowledge base is retrieved in parallel through a keyword matching module, a semantic vector matching module and a user portrait matching module, and the keyword matching module positions keyword segments through inverted indexes; according to the method, a three-way parallel recall engine for keyword matching, semantic vector matching and user portrait matching is constructed, multi-dimensional accurate retrieval of user questions is achieved, a keyword matching module adopts an inverted index to rapidly position keyword segments, and coverage of basic semantics is ensured; the semantic vector matching module calculates cosine similarity through a Sension-BERT model, and captures deep semantic association between a user question and a knowledge base fragment; the user portrait matching module generates a retrieval weight based on a user identity dynamic loading rule base.
Owner:CHONGQING UNIV OF EDUCATION

Text object indexing method, object storage system and related devices

This application relates to a text object indexing method, an object storage system and related devices. Among them, the method includes: after the front-end service module writes a text object into the key-value storage database system, it adds a write success message of the text object to the text write message queue; the text analysis scheduling module consumes the write success message in the text write message queue, generates a text analysis task for the text object and schedules the text analysis module to execute the text analysis task, obtains a set of representative keywords of the text object as the metadata of the text object and updates it to the key-value storage database system, and adds a metadata update message of the text object to the metadata update message queue; so that the index service module reads the metadata of the text object from the key-value storage database system according to the metadata update message of the text object, constructs an inverted index of the text object, and provides an index service. Through the present invention, complex condition retrieval of text objects in an object storage system is realized.
Owner:ALIBABA CLOUD COMPUTING CO LTD

A data security sharing method, system and device

The present invention discloses a data security sharing method, system and device, relating to the technical field of data security. The method includes: a doctor creates an EMR for a patient seeking medical treatment, encrypts the generated EMR using the patient's public key, and stores the EMR ciphertext on a cloud server; information such as keywords, hash values, and logs of the patient's EMR are stored on a blockchain, and retrieval services are provided using the keywords to return the hash value and index of the EMR. The blockchain ciphertext retrieval method based on inverted index is used to retrieve the EMR ciphertext on the cloud server; this method realizes efficient ciphertext retrieval, thereby comprehensively improving the deficiencies of the existing electronic medical record system in terms of data sharing, efficiency and security.
Owner:BEIJING INFORMATION SCI & TECH UNIV +1

Data retrieval method and device, equipment, storage medium and computer program product

The invention discloses a data retrieval method and device, equipment, a storage medium and a computer program product, and relates to the technical field of data retrieve.The method comprises the steps that word segmentation processing is conducted on long text data, and an inverted index corresponding to the long text data is generated according to a word segmentation result; converting a business rule contained in the expert rule model into a rule element, wherein the rule element represents a rule for searching the business rule in the long text data; and generating a queried statement according to the rule element, and retrieving the long text data based on the queried statement and the inverted index. According to the method, the word segmentation result of the long text data is structured through the inverted index, and meanwhile, the queriable statement is generated according to the rule element converted from the business rule, so that abstract expression of complex logic is realized to enhance manageability and reusability of the business rule; therefore, accurate retrieval can be carried out in massive long text data in combination with the inverted index and the queried statements.
Owner:CHINA MERCHANTS BANK

Method and system for quickly retrieving similar texts

The invention discloses a similar text quick retrieval method and system, and belongs to the technical field of natural language processing and information retrieval, and the similar text quick retrieval method comprises the following specific steps: step 1, vectorizing semantics, and storing the vectorized semantics into a database; using a pre-training language model to respectively encode the query text and the documents in the text library into low-dimensional semantic vectors with fixed dimensions; step 2, hierarchical index retrieval: N candidate documents closest to the query vector are retrieved through a first-level index structure to form a first candidate set; according to the method, millisecond coarse screening is realized through parallel design of the HNSW graph index and the semantic cluster inverted index, the semantic cluster inverted index reduces a search space to a local optimal domain, global precision loss is avoided, a newly added text only needs one-time cluster center distance calculation, and index updating complexity is greatly reduced.
Owner:SHANGHAI DONGYONG NETWORK TECH CO LTD

Commercial bank blacklist management method based on big data platform and Elasticsearch

The invention discloses a commercial bank blacklist management method based on a big data platform and Elasticsearch, and the method comprises the steps: building a blacklist data center based on the big data platform, integrating multi-source heterogeneous data, and carrying out the real-time data cleaning and standardization processing; establishing a distributed real-time retrieval engine by utilizing Elasticsearch, and constructing a reverse index for the processed data; establishing an associated risk map model, dynamically analyzing the association degree between the client and the blacklist main body through a map calculation engine, generating a risk conduction coefficient, and storing the risk conduction coefficient; during business handling, a multi-condition combination query request is initiated to Elasticsearch through an ESB (Enterprise Service Bus) real-time calling interface; a dynamic updating mechanism is adopted, and the blacklist state is automatically updated when a risk threshold value is triggered. According to the method, by integrating multi-source heterogeneous data, millisecond risk interception is realized, a dynamic association graph is constructed, the active defense capability of a bank on risks such as fraudulent transactions, credit default and money laundering behaviors is improved, and the method is suitable for risk management and control of core business scenes such as pre-loan auditing, transaction monitoring and anti-money laundering.
Owner:BANK OF GUIYANG CO LTD

Geographic information resource governance system and method in localized cloud environment

The invention discloses a geographic information resource governance system and method in a localized cloud environment, a multi-mode geographic information resource connection convergence module provides a dynamically expanded geographic information resource connection capability, and provides real-time calculation, offline calculation and memory calculation capabilities in access and integration processes at the same time; the semantic enhancement geographic information resource efficient organization module is combined with a distributed key value data structure to generate an inverted index; meanwhile, the organization and storage modes of index data are dynamically adjusted in geographic information resource query and access operations, and heat perception dynamic partitioning is achieved; and the multi-granularity geographic information resource online service module divides the multi-modal geographic information resources from a resource level, a target level and an element level, predicts service types, interface standards and underlying environment requirements of geographic information resource release, and realizes rapid release of the multi-modal geographic information resources. According to the invention, effective treatment of multi-modal geographic information resources in a localized cloud environment can be realized.
Owner:SUZHOU AEROSPACE INFORMATION RES INST

Enhanced symmetric searchable encryption method and system

The invention provides an enhanced symmetric searchable encryption method and system, and relates to the field of data security and privacy protection. The method comprises the steps that a plaintext document set is symmetrically encrypted, keywords are extracted, and an encryption index token is generated through a pseudo-random function; constructing a reverse index structure, and mapping the token to a document identifier; dynamic insertion, deletion and modification of documents are supported, and indexes are synchronously updated in real time; cipher query is realized by generating a search token, and after the server returns a matching result, the client decrypts the matching result to obtain an original text. According to the system, through multi-keyword Boolean logic query, an index structure is compressed by adopting a Bloom filter, the retrieval efficiency is improved in combination with a jump pointer or a B + tree, a false query strategy is introduced to hide an access mode, and the anti-attack ability is enhanced. Compared with an existing scheme, the method has the advantages that the searchability, the system performance and the privacy protection capability of the encrypted data are improved, and the method is suitable for security data retrieval application in cloud computing, database systems and block chain environments.
Owner:浪潮智能终端有限公司

Searching method and device

The embodiment of the invention provides a search method and related equipment / products, and relates to the technical field of Internet. The search method comprises the following steps: receiving a search request sent by a client, wherein the search request carries an identifier of a target object; determining a target index entry from a plurality of index entries of an object reverse index based on the identifier of the target object; the object reverse index is constructed based on video contents of a plurality of live broadcasting rooms, the video contents correspond to one or more objects, and each index entry corresponds to one object and indicates an identifier of the live broadcasting room associated with the object; determining an identifier of a target live broadcast room associated with the target object according to the target index entry; and generating a target live broadcast room list according to the identifier of the target live broadcast room, and pushing the target live broadcast room list to the client. According to the technical scheme, a more visual, accurate and comprehensive search result can be provided, so that the search experience and the platform competitiveness are effectively improved.
Owner:SHANGHAI BILIBILI TECH CO LTD

Data multivariate storage and full-text retrieval system

The invention relates to the technical field of computer big data processing and retrieval, and discloses a data multivariate storage and full-text retrieval system which is characterized in that an isomorphic hash intake module is used for truncating an input data stream into logic data blocks, extracting lexical elements to generate isomorphic hash packets and distributing the isomorphic hash packets to a bottom layer for storage; the state monitoring module maintains heat potential energy and queries fingerprints, updates the state according to scanning feedback, and sends out a trigger signal when the heat exceeds a dimension rising threshold value; the differential projection module responds to the signal to read the isomorphic hash packet, and constructs a differential inverted index by calculating the query fingerprint and the intersection of the query fingerprint; and the cascade retrieval module executes memory search when the index is mounted, and scans the logic data block and feeds back a result to the state monitoring module if the index is not hit. According to the method, the heat potential energy of the data block is tracked through the state monitoring module, index construction is triggered only when query accumulation reaches the dimension raising threshold value, and cold start resource consumption of the system is obviously reduced.
Owner:SICHUAN LEWEI TECH CO LTD

System and method for generating SQL (Structured Query Language) by natural language based on ES, knowledge base and interaction enhancement

The invention relates to the technical field of language generation algorithms, and discloses a natural language generation SQL method based on ES, knowledge base and interaction enhancement, and the method comprises the following four sequentially executed core steps: data feature preprocessing: carrying out deep analysis (including field type, constraint relationship, data distribution and the like) on a database table structure and data features, and carrying out data feature preprocessing; and after data de-duplication is executed, the ES is imported, and an independent structured index is established according to a table-level isolation principle to provide basic data for subsequent retrieval. According to the method, the keyword matching efficiency and the semantic understanding depth are both considered through collaborative retrieval of the double-storage-layer knowledge base in combination with Elasticsearch keyword inverted index and Milvus vector library semantic retrieval. And the results are fused through the RRF algorithm, so that the retrieval precision is remarkably improved, the limitation of a single retrieval mode is solved, and the generated SQL better fits the database structure and the user intention.
Owner:QUALITY ENERGY BODY TECHNOLOGY (TIANJIN) CO LTD +1

Log management method and device, equipment and medium

The embodiment of the invention provides a log management method and device, equipment and a medium. The method comprises the steps that query keywords for logs are obtained; determining an association keyword according to the query keyword; the association keyword is generated according to semantic association of the query keyword; determining a reverse index field and a balanced tree index field from at least one field contained in the association keyword; determining at least one candidate log according to a reverse index field, a reverse index corresponding to the reverse index field, a balance tree index field and a balance tree index corresponding to the balance tree index field; respectively determining content keywords of at least one candidate log; according to the content keyword and the association keyword, respectively determining a display score of the at least one candidate log; and in the at least one candidate log, taking the log with the display score greater than a preset score threshold as a target log, and outputting the target log to the user, thereby improving the accuracy and efficiency of log query.
Owner:CHINA TELECOM CORP LTD

Data retrieval system of mixed storage and unified query layer based on AI cloud desktop

The invention provides a mixed storage and unified query layer data retrieval system based on an AI cloud desktop, which belongs to the technical field of AI intelligent application and comprises a unified query access layer, an SQL optimizer, a vectorization execution engine, a mixed storage engine and a metadata management center. Column storage, inverted index and row storage are organically combined, services are provided for the outside through a unified SQL query layer, and intelligent routing and vectorization computing technologies are introduced. The method is characterized in that an optimal storage engine is automatically selected for execution according to a query mode, the problems that a single inverted index architecture is low in efficiency and high in resource consumption under the scenes of ad-hoc analysis, high-concurrency point search and large-scale aggregation are solved, and better comprehensive query performance, lower data storage cost and higher system expansibility than traditional Elasticsearch are achieved.
Owner:INSPUR COMM TECH CO LTD

Large language model and multi-modal model retrieval method, electronic equipment and medium

The invention provides a retrieval method for a large language model and a multi-modal model, electronic equipment and a medium. The method comprises the steps of performing first preprocessing on a knowledge base to obtain a vector database; performing second preprocessing on the knowledge base to obtain a reverse index database; obtaining a search word; performing third preprocessing on the search word to obtain a first search word set, performing retrieval by a vector database according to the first search word set, and obtaining a vector recall slice set according to the vector database; performing fourth preprocessing on the search word to obtain a second search word set, performing retrieval by a reverse index database according to the second search word set, and recalling a text slice set according to feedback of the reverse index database; and receiving the vector recall slice set and the text slice set, combining the vector recall slice set and the text slice set into a recall slice set, processing the recall slice set, and inputting the processed recall slice set into a large language model to obtain a retrieval result. Therefore, the retrieval result provided by the invention is more accurate.
Owner:STATE GRID BUSINESS TRAVEL CLOUD TECH CO LTD