Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

234 results about "Inverted index" patented technology

In computer science, an inverted index (also referred to as a postings file or inverted file) is a database index storing a mapping from content, such as words or numbers, to its locations in a table, or in a document or a set of documents (named in contrast to a forward index, which maps from documents to content). The purpose of an inverted index is to allow fast full-text searches, at a cost of increased processing when a document is added to the database. The inverted file may be the database file itself, rather than its index. It is the most popular data structure used in document retrieval systems, used on a large scale for example in search engines. Additionally, several significant general-purpose mainframe-based database management systems have used inverted list architectures, including ADABAS, DATACOM/DB, and Model 204.

Intelligent search engine system and method based on NLP and vector hybrid retrieval

The invention relates to the technical field of natural language processing, in particular to an intelligent search engine system and method based on NLP and vector hybrid retrieval, and the system comprises a query analysis module which is used for receiving a query statement input by a user and obtaining structured query information based on the query statement; the symbiotic index module comprises a sparse extension index unit which is used for constructing a sparse inverted index table based on keywords in the structured query information; the vectorization hypergraph index unit is used for forming a hypergraph index based on the business data; the recall arrangement module is used for acquiring a candidate object set according to the structured query information; the constraint rearrangement module is used for calculating a joint priority score between query and candidate objects based on the candidate object set, and obtaining a sorting result under constraint conditions of meeting a supplier proportion, a category proportion and a price interval; and the evidence generation module is used for generating evidence information based on the sorting result and outputting the evidence information to the user interface.
Owner:BEIJING JINGNENG TENDERING & COLLECTIVE PROCUREMENT CENT CO LTD

Electronic archive intelligent retrieval method and system based on block chain

The invention discloses an electronic archive intelligent retrieval method and system based on a block chain, and relates to the technical field of archive retrieval, and the method comprises the steps: extracting text features of an electronic archive, and generating metadata; generating a content hash value from the electronic file original text; writing the metadata, the content hash value and the ABE access strategy into the block chain smart contract; constructing a reverse index based on the metadata semantic tag, wherein an index entry is associated with a block chain storage address; calculating a root hash and anchoring the root hash to the block chain; analyzing a keyword and a digital identity certificate in the user retrieval request; calling an intelligent contract to verify whether the user attribute accords with the ABE access strategy of the target file; retrieving an encrypted hash list matched with the file in the distributed index; acquiring the encrypted file fragments from the distributed storage system; verifying data integrity; and combining the fragments to generate a final retrieval result. The method has the advantages that through deep coupling of the block chain and attribute encryption, an electronic archive management system considering security and intelligent retrieval is constructed.
Owner:BEIJING RUIYUN ARCHIVES MANAGEMENT CO LTD

Method and system for retrieving DOCX document content based on keywords

The invention belongs to the technical field of text processing, and particularly relates to a method and system for retrieving DOCX document content based on keywords, which comprises the following steps: analyzing an Office Open XML structure of a DOCX document, combining with multi-dimensional features such as style names, and utilizing a title classification score model to accurately distinguish a title and a text, so that a semantic hierarchical structure of the document is effectively reserved; and secondly, a multi-level semantic extension mechanism is introduced, and a Sension-BERT, a HowNet knowledge base and a Word2Vec model are fused, so that intelligent extension of synonyms and synonyms of keywords is realized, and the recall rate and semantic understanding ability of retrieval are remarkably improved. And in addition, a BM25 model is combined with paragraph length normalization and structure position weight to calculate a correlation score, so that retrieval results are sorted more accurately and reasonably. The construction of the reverse index is combined with the position coding and compression optimization strategy, and the retrieval efficiency and the storage performance are both considered.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Retrieval method and device based on static word embedding, computer equipment and medium

The invention relates to a retrieval method and device based on static word embedding, computer equipment, a computer readable storage medium and a computer program product. The method comprises the steps of performing word segmentation on an original text to obtain a first word segmentation result, training a static word embedding model by utilizing the first word segmentation result to obtain word vectors, and generating a synonym word library; the method comprises the following steps: establishing a full-text inverted index by utilizing an original text, and expanding query words by utilizing a synonym library in a retrieval stage; encoding the original text into a semantic vector by using a semantic generation model, and constructing a vector index based on the semantic vector; based on a to-be-queried text in a user query request, performing retrieval by using the full-text inverted index to obtain a first candidate document, and performing retrieval by using the vector index to obtain a second candidate document; performing fusion processing on the first candidate document and the second candidate document to obtain a target candidate document; and inputting the target candidate document into the text generation model to obtain a retrieval result. By adopting the method, the accuracy of text retrieval can be improved.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Data segmentation and query processing system under distributed architecture

The invention relates to the technical field of computer data processing, and discloses a data segmentation and query processing system under a distributed architecture. The system comprises a data segmentation module which divides data blocks based on a self-adaptive partitioning algorithm; the distributed storage module is used for distributing data blocks by using a consistent Hash algorithm and dynamically adjusting the distribution of virtual nodes; the index construction module is used for generating a global index table through a multi-layer inverted index algorithm; the query optimization module is used for converting query statements by utilizing a query rewriting algorithm; and the result aggregation module is used for processing the sub-query results by adopting a parallel aggregation algorithm. The system can dynamically process data according to data characteristics, optimizes the data storage and query process, effectively improves the data processing efficiency and query performance of a distributed system, and is suitable for scenes such as big data storage and analysis.
Owner:ZHEJIANG JINGJING TECH CO LTD

Standard address pipeline data association management method and system based on multistage collaborative coding

The invention discloses a standard address pipeline data association management method and system based on multi-level collaborative coding, belongs to the technical field of government affair digitization, and aims to solve the technical problem of how to generate and manage standard address data in government affair business, realize automatic triggering and standardized circulation of cross-department business data and improve the management efficiency. According to the technical scheme, the method comprises the steps of cross-department business triggering: automatically triggering a road and courtyard address code generation process through construction project land pre-examination and construction project planning permission business events; a standard road name is generated based on road pre-examination data, a road unified identification code is synthesized according to a 4 + 6 + 9 segmentation rule, a house number is generated according to planning permission data, and a courtyard unified identification code is generated in combination with a naming system; the house prediction data triggers the generation of a house unified identification code; and data dynamic association: constructing a reverse index based on the address coding feature vector, and realizing millisecond-level accurate matching of the address and population information through a distributed query engine.
Owner:浪潮智慧城市科技有限公司

Hotel room type matching method and device

The invention discloses a hotel room type matching method and device, and the method comprises the steps: obtaining the original information of a newly-added hotel and room type, carrying out the preprocessing of the original information, and generating a standardized hotel and room type feature data set; constructing a multi-level feature vector based on the standardized feature data set; screening and generating a candidate matching set in combination with an inverted index and a locality sensitive hashing algorithm; a matching score matrix is calculated through fusion of the rule matching model, the machine learning model and the deep learning model; generating a final matching conclusion through Bayesian uncertainty estimation and threshold dynamic adjustment; and optimizing the matching model based on the matching conclusion and the business feedback data. Through multi-dimensional feature extraction, multi-model fusion matching and confidence evaluation, the problem of accuracy of hotel room type cross-supplier matching is solved, and the matching efficiency and robustness are improved.
Owner:GUIZHOU YOUTEYUN TECH CO LTD

Text content index automatic identification method based on semantics

The invention discloses a semantic-based text content index automatic identification method, which relates to the technical field of information retrieval, and comprises the following steps: initializing sparse projection and LSH signature, performing iterative optimization by using a Lagrange duality form and gradient update, adjusting hash digits, obtaining a fragment index through a k-d tree, and constructing an inverted index. According to the method, compression is performed through Delta coding, an index map is constructed based on Jaccard similarity, compression is performed through WebGraph, CSNMF is used in combination with Z-Laplacian regularization, a low-rank basis matrix and a low-rank coding matrix are generated, a compressed inverted index is reconstructed after iterative optimization, and reconstructed inverted index entries are generated. According to the method, through multi-resolution hash table initialization, joint feature optimization, local adaptive quantization and low-rank index reconstruction, the semantic expression ability and the compression effect of an index structure are improved, the index precision and efficiency are improved, and intelligent identification of index content is achieved.
Owner:BEIJING GEPU TECHNOLOGY CO LTD

Intelligent Agent collaborative decision-making system for environment monitoring

The invention relates to the technical field of water quality monitoring and analysis, in particular to an intelligent Agent collaborative decision-making system oriented to environment monitoring. Comprising the steps that original data including water quality parameters and environment data are collected in real time; performing noise elimination, missing value filling and standardization processing on the collected original data, and outputting preprocessed data; the method comprises the following steps: integrating multi-source data including historical monitoring data, pollution source information and treatment cases, performing vectorization coding on the multi-source data by adopting a BERT model, and constructing a retrieval architecture based on hierarchical indexes and inverted indexes of a hierarchical navigable small world to obtain a dynamic water quality knowledge base; and on the basis of a retrieval enhancement generation technology, associating the preprocessed data with the dynamic water quality knowledge base, predicting a water quality change trend through a time sequence prediction model, and generating an abnormality diagnosis report and a closed-loop treatment scheme. According to the invention, analysis and early warning of water quality are realized through collaborative decision-making of the intelligent Agent.
Owner:ZHEJIANG GUANGCHUAN ENG CONSULTING CO LTD

Intelligent dialogue knowledge base optimization method based on multi-path recall fusion and self-updating

The invention discloses an intelligent dialogue knowledge base optimization method based on multi-path recall fusion and self-updating, and relates to the technical field of artificial intelligence. The method comprises the following steps that A, a multi-path parallel recall engine is constructed, a knowledge base is retrieved in parallel through a keyword matching module, a semantic vector matching module and a user portrait matching module, and the keyword matching module positions keyword segments through inverted indexes; according to the method, a three-way parallel recall engine for keyword matching, semantic vector matching and user portrait matching is constructed, multi-dimensional accurate retrieval of user questions is achieved, a keyword matching module adopts an inverted index to rapidly position keyword segments, and coverage of basic semantics is ensured; the semantic vector matching module calculates cosine similarity through a Sension-BERT model, and captures deep semantic association between a user question and a knowledge base fragment; the user portrait matching module generates a retrieval weight based on a user identity dynamic loading rule base.
Owner:CHONGQING UNIV OF EDUCATION

Data retrieval method and device, equipment, storage medium and computer program product

The invention discloses a data retrieval method and device, equipment, a storage medium and a computer program product, and relates to the technical field of data retrieve.The method comprises the steps that word segmentation processing is conducted on long text data, and an inverted index corresponding to the long text data is generated according to a word segmentation result; converting a business rule contained in the expert rule model into a rule element, wherein the rule element represents a rule for searching the business rule in the long text data; and generating a queried statement according to the rule element, and retrieving the long text data based on the queried statement and the inverted index. According to the method, the word segmentation result of the long text data is structured through the inverted index, and meanwhile, the queriable statement is generated according to the rule element converted from the business rule, so that abstract expression of complex logic is realized to enhance manageability and reusability of the business rule; therefore, accurate retrieval can be carried out in massive long text data in combination with the inverted index and the queried statements.
Owner:CHINA MERCHANTS BANK

Method and system for quickly retrieving similar texts

The invention discloses a similar text quick retrieval method and system, and belongs to the technical field of natural language processing and information retrieval, and the similar text quick retrieval method comprises the following specific steps: step 1, vectorizing semantics, and storing the vectorized semantics into a database; using a pre-training language model to respectively encode the query text and the documents in the text library into low-dimensional semantic vectors with fixed dimensions; step 2, hierarchical index retrieval: N candidate documents closest to the query vector are retrieved through a first-level index structure to form a first candidate set; according to the method, millisecond coarse screening is realized through parallel design of the HNSW graph index and the semantic cluster inverted index, the semantic cluster inverted index reduces a search space to a local optimal domain, global precision loss is avoided, a newly added text only needs one-time cluster center distance calculation, and index updating complexity is greatly reduced.
Owner:SHANGHAI DONGYONG NETWORK TECH CO LTD

Commercial bank blacklist management method based on big data platform and Elasticsearch

The invention discloses a commercial bank blacklist management method based on a big data platform and Elasticsearch, and the method comprises the steps: building a blacklist data center based on the big data platform, integrating multi-source heterogeneous data, and carrying out the real-time data cleaning and standardization processing; establishing a distributed real-time retrieval engine by utilizing Elasticsearch, and constructing a reverse index for the processed data; establishing an associated risk map model, dynamically analyzing the association degree between the client and the blacklist main body through a map calculation engine, generating a risk conduction coefficient, and storing the risk conduction coefficient; during business handling, a multi-condition combination query request is initiated to Elasticsearch through an ESB (Enterprise Service Bus) real-time calling interface; a dynamic updating mechanism is adopted, and the blacklist state is automatically updated when a risk threshold value is triggered. According to the method, by integrating multi-source heterogeneous data, millisecond risk interception is realized, a dynamic association graph is constructed, the active defense capability of a bank on risks such as fraudulent transactions, credit default and money laundering behaviors is improved, and the method is suitable for risk management and control of core business scenes such as pre-loan auditing, transaction monitoring and anti-money laundering.
Owner:BANK OF GUIYANG CO LTD

Enhanced symmetric searchable encryption method and system

The invention provides an enhanced symmetric searchable encryption method and system, and relates to the field of data security and privacy protection. The method comprises the steps that a plaintext document set is symmetrically encrypted, keywords are extracted, and an encryption index token is generated through a pseudo-random function; constructing a reverse index structure, and mapping the token to a document identifier; dynamic insertion, deletion and modification of documents are supported, and indexes are synchronously updated in real time; cipher query is realized by generating a search token, and after the server returns a matching result, the client decrypts the matching result to obtain an original text. According to the system, through multi-keyword Boolean logic query, an index structure is compressed by adopting a Bloom filter, the retrieval efficiency is improved in combination with a jump pointer or a B + tree, a false query strategy is introduced to hide an access mode, and the anti-attack ability is enhanced. Compared with an existing scheme, the method has the advantages that the searchability, the system performance and the privacy protection capability of the encrypted data are improved, and the method is suitable for security data retrieval application in cloud computing, database systems and block chain environments.
Owner:浪潮智能终端有限公司

Data multivariate storage and full-text retrieval system

The invention relates to the technical field of computer big data processing and retrieval, and discloses a data multivariate storage and full-text retrieval system which is characterized in that an isomorphic hash intake module is used for truncating an input data stream into logic data blocks, extracting lexical elements to generate isomorphic hash packets and distributing the isomorphic hash packets to a bottom layer for storage; the state monitoring module maintains heat potential energy and queries fingerprints, updates the state according to scanning feedback, and sends out a trigger signal when the heat exceeds a dimension rising threshold value; the differential projection module responds to the signal to read the isomorphic hash packet, and constructs a differential inverted index by calculating the query fingerprint and the intersection of the query fingerprint; and the cascade retrieval module executes memory search when the index is mounted, and scans the logic data block and feeds back a result to the state monitoring module if the index is not hit. According to the method, the heat potential energy of the data block is tracked through the state monitoring module, index construction is triggered only when query accumulation reaches the dimension raising threshold value, and cold start resource consumption of the system is obviously reduced.
Owner:SICHUAN LEWEI TECH CO LTD

System and method for generating SQL (Structured Query Language) by natural language based on ES, knowledge base and interaction enhancement

The invention relates to the technical field of language generation algorithms, and discloses a natural language generation SQL method based on ES, knowledge base and interaction enhancement, and the method comprises the following four sequentially executed core steps: data feature preprocessing: carrying out deep analysis (including field type, constraint relationship, data distribution and the like) on a database table structure and data features, and carrying out data feature preprocessing; and after data de-duplication is executed, the ES is imported, and an independent structured index is established according to a table-level isolation principle to provide basic data for subsequent retrieval. According to the method, the keyword matching efficiency and the semantic understanding depth are both considered through collaborative retrieval of the double-storage-layer knowledge base in combination with Elasticsearch keyword inverted index and Milvus vector library semantic retrieval. And the results are fused through the RRF algorithm, so that the retrieval precision is remarkably improved, the limitation of a single retrieval mode is solved, and the generated SQL better fits the database structure and the user intention.
Owner:QUALITY ENERGY BODY TECHNOLOGY (TIANJIN) CO LTD +1

Log management method and device, equipment and medium

The embodiment of the invention provides a log management method and device, equipment and a medium. The method comprises the steps that query keywords for logs are obtained; determining an association keyword according to the query keyword; the association keyword is generated according to semantic association of the query keyword; determining a reverse index field and a balanced tree index field from at least one field contained in the association keyword; determining at least one candidate log according to a reverse index field, a reverse index corresponding to the reverse index field, a balance tree index field and a balance tree index corresponding to the balance tree index field; respectively determining content keywords of at least one candidate log; according to the content keyword and the association keyword, respectively determining a display score of the at least one candidate log; and in the at least one candidate log, taking the log with the display score greater than a preset score threshold as a target log, and outputting the target log to the user, thereby improving the accuracy and efficiency of log query.
Owner:CHINA TELECOM CORP LTD

Data retrieval system of mixed storage and unified query layer based on AI cloud desktop

The invention provides a mixed storage and unified query layer data retrieval system based on an AI cloud desktop, which belongs to the technical field of AI intelligent application and comprises a unified query access layer, an SQL optimizer, a vectorization execution engine, a mixed storage engine and a metadata management center. Column storage, inverted index and row storage are organically combined, services are provided for the outside through a unified SQL query layer, and intelligent routing and vectorization computing technologies are introduced. The method is characterized in that an optimal storage engine is automatically selected for execution according to a query mode, the problems that a single inverted index architecture is low in efficiency and high in resource consumption under the scenes of ad-hoc analysis, high-concurrency point search and large-scale aggregation are solved, and better comprehensive query performance, lower data storage cost and higher system expansibility than traditional Elasticsearch are achieved.
Owner:INSPUR COMM TECH CO LTD

Book resource data security collection retrieval monitoring system based on big data

The invention discloses a book resource data security collection retrieval monitoring system based on big data, and relates to the technical field of information technology and data security, and the system comprises a permission-aware association reasoning retrieval module which is in communication connection with a knowledge graph and dynamic index construction module and is used for receiving a user retrieval request and sending the user retrieval request to a database; analyzing a retrieval intention and obtaining a real-time permission context of a user to determine a security level threshold value, verifying a security attribute voucher of the data when retrieving the mixed index, only putting the data with the security level not higher than the threshold value into a result set, and performing association reasoning based on the knowledge graph under permission constraint to expand a result. According to the method, the mixed index structure fusing the inverted index, the vector index and the graph index is constructed, the security attribute voucher is associated, efficient full-text retrieval, semantic retrieval and association reasoning are supported, meanwhile, it is ensured that the retrieval process and result are strictly constrained by permission, and maximum mining of data values on the premise of security is achieved.
Owner:GUANGDONG POLYTECHNIC OF IND & COMMERCE

Multi-keyword private information retrieval method and system based on inverted index

The invention provides a multi-keyword private information retrieval method and system based on inverted indexes, and belongs to the field of information security and privacy computing. The method comprises the following steps that: a server splits an original database into a reverse index part of a keyword-index set and a key value pair part of an index-data value; mapping the index set into an integer by utilizing Godel coding; processing the data by adopting probability batch coding and binary random linear coding and publishing a hash function; the client maps a multi-keyword query into a bucket by using a hash function to generate a query vector, and the query vector is sent to the server after being subjected to fully homomorphic encryption; the server performs homomorphic calculation on a matching result in a ciphertext state and returns the result; and the client obtains a matching index set after decryption verification, and then initiates a second round of query to obtain final data. According to the method, symmetric privacy protection of multi-keyword non-primary key query in a single-server environment is realized, leakage of query content and database information is effectively prevented, and private information retrieval security is improved.
Owner:UNIV OF JINAN

Intelligent office knowledge system and dual-mode precise retrieval method

The invention relates to the technical field of data processing, and discloses an intelligent office knowledge system and a dual-mode precise retrieval method.The method comprises the steps that firstly, a keyword retrieval mode is adopted, a text needing to be retrieved is preprocessed, specifically, a self-defined dictionary is loaded through Jiaba word segmentation Chinese or NLTK English, stop words are deleted and filtered out, stem extraction is conducted, and inverted index construction is conducted; mapping is established for the keywords; according to the semantic retrieval mode, firstly, text vectorization is carried out, and a document is converted into a vector by using Word2Vec or BERT; similarity calculation is carried out, text semantics are converted into a vector Vquery, and cosine similarity is calculated; the result is weighted, the keyword retrieval score is equal to TF-IDF, the semantic retrieval score is equal to cos theta, and the fusion score is calculated; and the retrieved article sequence is eliminated again according to the fusion score. According to the method, the retrieval efficiency is improved, through the dual-mode accurate retrieval method, the user can find the needed information more quickly, the screening time is shortened, the working efficiency is improved, and the retrieval accuracy is improved.
Owner:GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD

Verifiable ciphertext retrieval method and device, equipment, medium and product

The invention discloses a verifiable ciphertext retrieval method and device, equipment, a medium and a product, and relates to the field of data security. The method comprises the steps that a server receives a query keyword set of a client; the server calculates a main keyword index address according to the main keyword of the query keyword set and searches the encrypted inverted index according to the main keyword index address; the server obtains candidate documents according to the encrypted document operation content in the encrypted inverted index; the client extracts a target data block position set based on a vector commitment tree according to the candidate document; the server generates a brief proof according to the target data block position set and determines a multi-keyword tag according to secondary keywords in the query keyword set; the server determines candidate results by using a label set according to the multi-keyword labels; and the client determines the document set meeting all the keywords according to the candidate result and the brief proof and then performs consistency check to obtain the retrieval result, and the method and the device reduce the calculation consumption while meeting the forward security.
Owner:GUIZHOU UNIV

Intelligent customer service knowledge base automatic updating method based on large language model

The invention discloses an intelligent customer service knowledge base automatic updating method based on a large language model, and relates to the technical field of knowledge base management, and the method comprises the steps: collecting a multi-source business document and session data, cleaning and merging the multi-source business document and the session data, calling the large language model to extract candidate entries, and binding evidence fingerprints; performing consistency comparison and change type judgment on the candidate entries and knowledge base entries, and generating and solidifying a difference abstract; generating a semantic patch according to the difference abstract, writing the semantic patch into a shadow partition, completing consistency verification, enabling atomic switching to take effect, synchronously updating a reverse index view and a vector index view of the intelligent customer service knowledge base, and recording a version mark; and performing online verification on a session data statistical result according to the version mark, rolling back to a state before the version mark when the verification result is abnormal, and confirming the semantic patch and continuously acting on the intelligent customer service knowledge base when the verification result is normal. The release risk is reduced; and the knowledge reliability and auditing performance are improved.
Owner:昆明双淼科技有限公司

A subgraph matching based graph similarity search method

The application discloses a graph similarity search method based on subgraph matching, divides data graphs into a plurality of mutually exclusive half-edge graphs, constructs an inverted index by taking subgraph embedding as a key and a data graph set containing the subgraph as a value, constructs a coordinate index based on size information of the data graph, realizes hierarchical filtering by using a mapping relationship between the two indexes, processes the half-edge by using a self-loop or loop-free strategy, designs a sequential embedding model, so that embeddings of two graphs with subgraph isomorphism have the same partial order position in a high-dimensional space, inputs a query graph, screens out a preliminary candidate set through the coordinate index, further screens the candidate set through the inverted index, accurately calculates a graph edit distance between the query graph and the candidate graph, and obtains a final result set, which not only avoids traditional subgraph isomorphism testing, but also can generate subgraph embedding in an offline stage, greatly shortens inference time of the model.
Owner:NANJING UNIV OF POSTS & TELECOMM

Intelligent question answering system-oriented method for identifying focus words in natural language question sentences

The invention discloses a method for identifying focus words in a natural language question for an intelligent question answering system, and relates to the technical field of natural language question answering. The invention provides a method for identifying focus words in natural language questions, so that a question answering system can more accurately understand concern points of a user; a prefix tree structure dominated by decision items is provided, an algorithm for mining strong focus association rules to identify focus words is further introduced based on the prefix tree, and the algorithm is more efficient than a classical association rule mining algorithm Apriori; an inverted index for a strong focus association rule is provided, and an algorithm for identifying focus words is further introduced based on the inverted index and is more efficient than sequential search; and a focus item set, a frequent focus item set, a focus association rule and a strong focus association rule are defined so as to better express information related to focus words.
Owner:YANGTZE NORMAL UNIVERSITY

Master data integration methods, apparatus, electronic devices and storage media

This application discloses a master data integration method, apparatus, electronic device, and storage medium. The method includes: acquiring multi-source heterogeneous master data; standardizing and cleaning the multi-source heterogeneous master data to generate multiple block keys and inverted index lists, and constructing a candidate comparison set; extracting and normalizing the multi-level similarity of each pair of candidate records in the candidate comparison set, and constructing a standardized feature vector; inputting the standardized feature vector into a learning model to obtain the matching probability of each pair of candidate records, and making hierarchical decisions based on the matching probability and at least one matching threshold; constructing an undirected matching graph and generating entity clusters from the undirected matching graph; determining the data quality score of each field of each candidate record within the entity cluster, determining the optimal value of each field based on the data quality score, and generating the integrated master data. This application significantly reduces the number of comparisons, improves comparison efficiency, eliminates the need for manually writing matching rules, and can adapt to various data types and new data sources.
Owner:ZHENHUA ZHIZAO (XIAN) TECH CO LTD

White list generation method and related device

The embodiment of the invention provides a white list generation method and a related device, and relates to the technical field of big data, and the method comprises the steps: carrying out the semantic recognition of a target policy, and determining first declaration data included in the target policy; obtaining a knowledge graph corresponding to the target policy, processing the first declaration data based on the knowledge graph, and obtaining second declaration data implied by the target policy; determining a declaration condition corresponding to the target policy based on the first declaration data and the second declaration data; acquiring first enterprise data declared by an enterprise; acquiring second enterprise data from an external data partner system; performing credible verification on the first enterprise data according to the second enterprise data, and determining target enterprise data; constructing a reverse index of the target enterprise data, and constructing a query statement of the target enterprise data according to the declaration condition; and according to the query statement and the inverted index, generating an enterprise white list meeting the declaration condition. According to the method, the generation efficiency and accuracy of the white list can be effectively improved.
Owner:CHINA CONSTRUCTION BANK +1

A sentiment analysis method based on retrieval and contrastive learning

The application provides a kind of sentiment analysis method based on retrieval and comparison learning in the technical field of natural language processing, comprising: step S10, obtaining a large amount of sentiment text data, and pre-processing each sentiment text data;Step S20, extract the entity in each pre-processed sentiment text data, label each entity to construct a sample, and then generate a sentiment dataset;Step S30, by Elasticsearch, the samples in the sentiment dataset are inverted indexed, and similar samples are retrieved for each sample;Step S40, create a sentiment classification model based on neural network, train the sentiment classification model using the sentiment dataset, and at the same time, use the contrast learning technology to narrow the vector distance between each sample and similar sample;Step S50, use the trained sentiment classification model to perform sentiment analysis.The application has the advantages of greatly improving the model sentiment representation ability, and greatly improving the sentiment classification performance.
Owner:XIAMEN UNIV

Graph anchor-based multi-agent meta-analysis literature extraction method and system

This invention discloses a method and system for extracting meta-analysis literature based on graph-anchored multi-agent systems. The method includes reconstructing unstructured PDFs into a hierarchical semantic document structure through visual document analysis and identifying functional area labels; applying functional area anchoring constraints to limit information search to specified functional area nodes; scheduling a multi-agent collaborative extraction network consisting of a reconnaissance agent, a logic assembly agent, and an indicator retrieval agent; and completing experimental variable identification, logical mapping between treatment and control groups, and cross-paragraph indicator retrieval through structured message dialogue and multi-round collaboration; employing a neural symbolic hybrid engine, where semantics are parsed by a large language model and precise mathematical aggregation operations and dimensional conversions are performed by a deterministic symbolic computation engine; and generating a structured data matrix with inverted index pointers. This application achieves high-precision, zero-error, and traceable automated extraction of meta-analysis literature data.
Owner:INNER MONGOLIA UNIVERSITY +1

Progressive question generation method based on semantic analysis and knowledge graph

The application discloses a progressive question generation method based on semantic analysis and a knowledge graph. The question generation demand text is preprocessed and semantically analyzed, the core knowledge point corresponding essential keyword set and the additional condition corresponding optional limiting word set are extracted, the concept mapping is carried out in the college professional knowledge graph, the curriculum outline is associated, the hierarchical semantic constraint framework containing hard constraint and soft constraint is constructed and consistency detection is carried out, the candidate materials are obtained by using the inverted index hard matching recall and the knowledge graph soft expansion recall, the candidate questions are generated by inputting the field fine-tuning large language model, the final questions are output by carrying out multi-objective learning sorting and reordering, teaching logic checking and quality evaluation, the incremental samples are formed by receiving user feedback, the generation model and the sorting model are updated, the accuracy, diversity and controllability of the question generation are improved, and closed-loop optimization is supported.
Owner:HOHAI UNIV