Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

717 results about "Semantic search" patented technology

Semantic search denotes search with meaning, as distinguished from lexical search where the search engine looks for literal matches of the query words or variants of them, without understanding the overall meaning of the query. Semantic search seeks to improve search accuracy by understanding the searcher's intent and the contextual meaning of terms as they appear in the searchable dataspace, whether on the Web or within a closed system, to generate more relevant results. Semantic search systems consider various points including context of search, location, intent, variation of words, synonyms, generalized and specialized queries, concept matching and natural language queries to provide relevant search results.

Multi-modal knowledge graph construction method and device based on large model and program product

The invention discloses a multi-modal knowledge graph construction method and device based on a large model and a program product, belongs to the technical field of artificial intelligence and knowledge engineering crossing, and particularly relates to a knowledge graph dynamic construction and evolution method based on a large language model technology and multi-modal data processing. The problems of symbol grounding and semantic understanding obstacle in the prior art are solved. According to the method, innovation and breakthrough are realized in three dimensions of knowledge acquisition, representation and reasoning by fusing deep learning and knowledge engineering technologies. The multi-modal knowledge graph construction method and device based on the large model and the program product are applied to the field of multi-modal knowledge graph construction and are suitable for specific task scenes such as intelligent question and answer, decision support and semantic search.
Owner:HARBIN INST OF TECH

Data reordering retrieval method and system based on RAG

The invention provides a data reordering retrieval method and system based on RAG, and the method comprises the steps: carrying out the dynamic semantic partitioning processing of an original document, generating corresponding document blocks, storing the document blocks in a vector database, constructing a hierarchical index, analyzing a received query request, extracting keywords in the query request, and carrying out the retrieval of the query request. Boolean keyword matching is carried out through an inverted index in the hierarchical index, semantic retrieval of keywords is carried out in a vector database, an initial candidate set is generated after multi-source retrieval results are fused, a query request and the initial candidate set are input into a generative reordering model, and a query result is obtained; according to the method, the initial candidate set is selected, the corresponding score is generated through the cross-modal attention mechanism in the generative reordering model, the initial candidate set is reordered according to the score, the final retrieval result is obtained, and the reordered document blocks are displayed, so that the overall retrieval speed is increased, the limitation of a single retrieval mode is avoided, and the accuracy of the retrieval result is improved.
Owner:TAIJI COMPUTER CORPORATION LIMITED

Low-altitude intelligent question and answer construction method and system based on dynamic parameters

The invention relates to a low-altitude intelligent question and answer construction method and system based on dynamic parameters. The method comprises the following steps: collecting low-altitude domain data, cleaning the low-altitude domain data, generating a semantic vector index, and constructing a low-altitude domain knowledge base based on the semantic vector index; receiving a natural language query of a user, analyzing a query intention, extracting keywords in the natural language query, and matching a corresponding candidate word quantity based on query types of the natural language query of the user, the query types at least comprising high-frequency phrase query and low-frequency long-tail query; and respectively carrying out fusion semantic retrieval and keyword retrieval, carrying out secondary sorting on the candidate results based on a preset resorter, preferentially sorting the candidate results related to the query intention, and outputting the corresponding candidate results. By adopting the method, a combined domain retrieval enhancement generation mechanism is provided, so that the professionality and accuracy of answers are improved; and the retrieved knowledge base content is re-screened to increase the hit probability of the knowledge base.
Owner:CHINA TELECOM UNMANNED TECHNOLOGY (JIANGSU) CO LTD

Multi-path recall retrieval method and system based on dynamic weight distribution and storage medium

The invention discloses a multi-path recall mixed retrieval method and system based on intelligent dynamic weight distribution, and aims to solve the problems that semantic comprehension and keyword matching are difficult to balance and the adaptability is poor due to the adoption of a fixed weight in the existing retrieval technology. The invention provides a multi-path recall mechanism fusing vector semantic retrieval, BM25 keyword retrieval and entity retrieval. A query feature vector containing 13-dimensional features such as semantic complexity, keyword density and entity coverage rate is constructed, a query type is recognized in combination with an SVM and a random forest integration model, a dynamic weight distribution algorithm is designed, and the final weight of each retrieval path is calculated in real time. And an adaptive multi-source enhanced reciprocal ranking fusion (AMSE-RRF) algorithm is further adopted to carry out optimization fusion on multiple paths of results, and a depth reordering model can be selected to improve the precision. According to the method, the accuracy and robustness of retrieval can be remarkably improved in multiple scenes of medical treatment, finance, government affairs and the like according to a millisecond-level self-adaptive adjustment strategy of query features.
Owner:DACE INFORMATION TECH CO LTD

Augmenting semantic search scores based on relevancy and popularity

Systems and methods for generating augmented search results are disclosed. An example method is performed by one or more processors of a search results ranking system and includes receiving a transmission over a communications network from a computing device associated with a user of the search results ranking system, the transmission including a search query, submitting, to a vector database, a token query matching a tokenized version of the search query against a plurality of data assets, submitting, to the vector database, one or more vector queries matching a vectorized version of the search query against the plurality of data assets, identifying, based on results of the token query and the one or more vector queries, contextually relevant results among the plurality of data assets, and generating augmented search results for the search query based on the contextually relevant results.
Owner:INTUIT INC

Rich-Media Document Auxiliary Generation Apparatus

Disclosed in the present disclosure is a rich-media document auxiliary generation apparatus. The apparatus comprises a material extraction module, a theme sorting module, a semantic retrieval module, a structured data text generation module, an illustration recommendation module and a video composition module. The present disclosure uses intelligent means to assist a user to efficiently generate a high-quality rich-media composite document, thereby quickly and accurately describing a theme event in an all-round way.
Owner:10TH RES INST OF CETC

Methods and apparatus for a retrieval augmented generative (RAG) artificial intelligence (AI) system

A non-transitory, processor-readable medium storing instructions that when executed by a processor, cause the processor to receive data artifacts, encode the artifacts to a standard data type, and compute, for each artifact, a hash function. The hash functions and encoded documents are stored in a first database. The processor is caused to tokenize the encoded artifacts, to produce tokens associated with natural-language identifiers extracted from the encoded artifacts. The processor is caused to transform, using an embedding model, the tokens to produce vectors that are stored in a second database and classified based on categories. The second database is configured to be queried to perform a semantic search in response to receiving a request from a user operating a user compute device. The processor is caused to retrieve, from the semantic search, a subset of vectors from the second database to be displayed on the user compute device.
Owner:FEDDATA HOLDINGS LLC

Power transformer fault auxiliary decision-making method and system based on knowledge graph and large language model

The invention discloses a power transformer fault auxiliary decision-making method and system based on a knowledge graph and a large language model, and the method comprises the steps: obtaining power transformer fault text data, carrying out the preprocessing, obtaining fault-related text and table data, preliminarily defining an ontology, selecting a part of text data for marking, and carrying out the recognition of the ontology; a plurality of named entity recognition and relation extraction models are trained, an optimal model is determined, then triple extraction is carried out on unlabeled text data, and a knowledge graph is constructed; predicting a potential entity relationship in the knowledge graph based on a link prediction model, and complementing the knowledge graph under the large language model and human assistance; constructing a semantic search model training data set based on the large language model and training a semantic search model; based on the complemented knowledge graph, the trained semantic search model is utilized to retrieve knowledge graph sub-graphs according to the problem, and auxiliary decision making is completed. According to the method, the problem of low potential knowledge utilization level in the existing power transformer fault auxiliary decision-making based on the knowledge graph is solved.
Owner:NANJING INST OF TECH

Government affair file information extraction and question and answer method and device and medium

The invention relates to a government affair file information extraction and question answering method and device and a medium, and the method comprises the steps: carrying out the entity extraction of a government affair file through employing a BERT-CRF joint model, and obtaining a structured entity set; performing relation extraction on the structured entity set to generate a semantic relation set between the entities; constructing a knowledge graph according to the structured entity set and the semantic relationship set, storing entity nodes into a graph database, and storing an embedded vector of an entity text into a vector database; when a query request of a user is received, relation query of the graph database and semantic retrieval of the vector database are carried out, sub-graph structures and semantic matching vectors related to query are extracted, and a mixed retrieval result is obtained; and inputting the mixed retrieval result into a large language model, and generating a question and answer response text conforming to a preset format by applying a dynamic prompt template. According to the method, the document processing efficiency and accuracy are effectively improved, and a solid technical support is provided for intelligent management of government affair documents.
Owner:EVALUATION & DEMONSTRATION RES CENT OF THE CHINESE PEOPLES LIBERATION ARMY ACAD OF MILITARY SCI

Mixture language professional question and answer method based on mixed retrieval and retrieval enhancement generation

The invention provides a minority language professional question and answer method based on mixed retrieval and retrieval enhancement generation, which comprises the following steps: S1, constructing a multi-language knowledge constructing a vector index database and a term knowledge graph by using a multi-language model according to a related minority language document; s2, multi-layer mixed retrieval: cross-language document recall is realized through a multi-layer mixed retrieval module, and the multi-layer mixed retrieval module is composed of keyword retrieval, semantic retrieval and vector retrieval; s3, answer generation: performing answer generation through an adaptive multi-language model by using a retrieval enhancement generation module, and introducing rule constraint decoding and a dynamic attention mechanism in the generation stage to improve the professionality and accuracy of the answer; and S4, self-adaptive optimization and knowledge updating: through a user feedback reinforcement learning module, optimizing the model based on user error correction data and supervising updating of the knowledge base. According to the method, professional questions and answers of the minority language can be realized based on a small amount of professional data of the minority language, the recall rate of the document of the minority language is improved, and the generation quality is improved.
Owner:中关村视听产业技术创新联盟

Generative artificial intelligence (AI) construction specification interface

A method and system provide the ability to process a construction domain query. A natural language user query is obtained within a construction software system. The user query is pre-processed to validate the query. Text from the query is embed into search vectors for a semantic search. A data source having multiple different sections is obtained. The semantic search is performed within each of the sections and identifies semantically relevant sections. The relevant sections are consolidated into a contextual data prompt that is input into an LLM. The LLM, which is trained based on construction data, generates a response that identifies the relevant sections. The response and an identification of the relevant sections is output.
Owner:AUTODESK INC

Domain intelligent question-answering method and system based on multi-modal knowledge graph and RAG

The invention relates to the technical field of intelligent questioning and answering, in particular to a domain intelligent questioning and answering method and system based on a multi-modal knowledge graph and RAG, and the method comprises the steps: constructing a concept layer knowledge graph based on a directory structure of a domain multi-modal document, and constructing an instance layer knowledge graph based on document content; obtaining a user question, pruning and positioning the user question in combination with the concept layer knowledge graph and the thinking chain, and determining a target chapter; splitting the question into sub-questions through intention analysis, and performing semantic retrieval in the instance layer knowledge graph corresponding to the target chapter to obtain a graph retrieval result; optimizing the original problem based on the atlas retrieval result, and executing semantic retrieval in a vector database to obtain a vector retrieval result; and fusing the atlas retrieval result and the vector retrieval result to generate a preliminary answer, and performing iterative optimization until a final answer is generated. According to the method, the semantic coverage, the expression accuracy and the response efficiency of the vertical domain question-answering system are remarkably improved by constructing the multi-modal knowledge graph and optimizing the retrieval process.
Owner:HENAN UNIVERSITY

Knowledge construction method and system based on large model and RAG technology

The invention discloses a knowledge construction method and system based on a large model and an RAG technology, and belongs to the technical field of large models.According to the knowledge construction method and system based on the large model and the RAG technology, through cross-text-block mutual information calculation and RAG enhanced reasoning, the system can quantify statistical correlation between entities, and in combination with context information of an external knowledge base, potential incidence relations of cross-paragraphs or documents are mined, so that the knowledge construction efficiency is improved. A dynamic semantic segmentation strategy is adopted to ensure that semantics in text blocks are consistent, context continuity is maintained by retaining overlapped parts, multiple expressions of the same entity are comprehensively judged through a multi-dimensional anaphora resolution mechanism, the error resolution rate is reduced, dynamic segmentation and mixed retrieval are combined, and the problem of large model input length limitation is solved. And through semantic retrieval and keyword matching complementation, the recall rate is improved, and the key problems of semantic fracture, cross-text association missing, low entity alignment precision and the like in traditional knowledge construction are remarkably solved.
Owner:COMMUNICATION UNIVERSITY OF CHINA

Literature semantic search method and system based on elastic search

The invention discloses a literature semantic search method and system based on elastic search, and relates to data retrieval. The literature semantic search method comprises the steps that vectorization processing is conducted on a noun phrase list by means of a text2vec-based-multilingual model trained based on a CoSENT method; according to the semantic vector, performing approximate nearest neighbor search in a second retrieval module to obtain first candidate data; inputting the query text data into a first retrieval module, and performing keyword matching through a BM25 algorithm to obtain second candidate data; fusing the first candidate data and the second candidate data to obtain third candidate data; a Sequence Matcher algorithm is adopted to calculate character string similarity between expansion words in the third candidate data, a similarity threshold value is set based on the length of the longest common subsequence, duplicate removal is carried out, and fourth candidate data is obtained; and performing weight distribution based on positions and similarity scores on the fourth candidate data, and enhancing the distinction degree of the extension words by expanding a score interval to obtain extension word recommendation list data. According to the method, the accuracy of document retrieval is remarkably improved.
Owner:CHINA EDUCATIONAL PUBLICATIONS IMPORT & EXPORT CORP LTD

Knowledge question and answer rapid processing method and system based on artificial intelligence

The invention provides a knowledge question and answer rapid processing method and system based on artificial intelligence, and relates to the field of artificial intelligence. A multi-modal knowledge graph is constructed, collected multi-source teaching data is fused through a mixed retrieval strategy, and the mixed retrieval strategy comprises semantic retrieval, vector retrieval and metadata retrieval; multi-level question and answer processing is executed based on an RAG enhancement framework, a multi-modal input intention is analyzed, cross-library joint retrieval is performed, and an optimization answer is generated in combination with a teaching scene; distilling the global model to a lightweight TinyBERT architecture, dynamically optimizing question and answer quality through a cognitive reinforcement learning framework, positioning a key document from a comprehensive retrieval list, evaluating an optimized answer, and reconstructing an answer with a key document verification score; according to the invention, the professional skill level of teachers and students in the fields of artificial intelligence and large model application can be improved, and the personalized requirements of teachers and students in teaching, scientific research and innovation courses can be met.
Owner:RONGKE LIANCHUANG (TIANJIN) INFORMATION TECH CO LTD

Multi-modal heterogeneous knowledge fusion construction and semantic enhancement retrieval system based on large model

The invention relates to the technical field of multi-modal data processing and semantic retrieval, in particular to a multi-modal heterogeneous knowledge fusion construction and semantic enhancement retrieval system based on a large model, which comprises a data acquisition module, a semantic analysis module, a knowledge fusion module and a retrieval optimization module. Multi-modal data such as texts, images and audios are uniformly expressed and deeply analyzed by introducing a large model technology, a knowledge graph is dynamically constructed, a structure is optimized in combination with a user query intention, and meanwhile accurate sorting and screening are achieved through a semantic enhancement algorithm. According to the method, the semantic comprehension capability and the intelligent level of the system can be improved, the real-time and diversified scene requirements are met, and the accuracy and the adaptability of a retrieval result are remarkably enhanced.
Owner:ZHONGYU SOFTCOM (CHONGQING) INFORMATION TECH CO LTD

Heterogeneous network resource virtualization modeling and intelligent arrangement method and system

The invention provides a heterogeneous network resource virtualization modeling and intelligent arrangement method and system based on a knowledge graph, and relates to unified modeling, dynamic retrieval and intelligent resource arrangement of heterogeneous network equipment. The method specifically comprises: 1, a unified modeling method based on a knowledge graph: integrating protocol attributes, dynamic states and topological relationships of heterogeneous devices such as a 5G base station, an SDN switch, a router, an Internet of Things gateway and the like into a structured knowledge graph, breaking the barrier of a manufacturer private data model, and realizing semantic-level collaborative scheduling of cross-domain resources; 2, designing a semantic retrieval engine: querying dynamic conditional reasoning through a natural language, replacing traditional manual rule definition, and improving retrieval response speed and accuracy; and 3, developing a graph-driven intelligent arrangement framework: combining a graph neural network (GNN) and reinforcement learning (RL), automatically generating a resource allocation strategy according to a real-time network state, reducing manual intervention and improving the resource utilization rate.
Owner:NO 50 RES INST OF CHINA ELECTRONICS TECH GRP

Intelligent log retrieval and analysis system based on model context protocol (MCP)

The invention discloses an intelligent log retrieval and analysis system based on a model context protocol MCP, and relates to the technical field of log analysis. The system comprises an MCP log semantic conversion and retrieval engine, a log format irrelevant feature extraction and indexer, a dynamic log format recognition and context enhancement system, a real-time retrieval analysis and aggregation controller and a multi-level semantic retrieval and visualization engine. According to the context enhancement system, unknown formats can be identified, semantic tags can be complemented, and semantic consistency and behavior tracking capability can be improved. The retrieval analysis controller supports semantic expression analysis, strategy generation and feedback closed loop, and strategy scheduling and multi-dimensional aggregation analysis are achieved. And finally, outputting a structured semantic result by the system, and visually displaying the structured semantic result through components such as a semantic composition device and a context expander. All the modules are managed in a unified mode through a capability registration mechanism, dynamic arrangement and upstream and downstream closed-loop linkage are supported, and a log intelligent analysis framework with high semantic driving and a clear structure is formed.
Owner:SHANGHAI NETIS TECH CO LTD

Distributed storage method based on source code semantic partitioning

The invention provides a distributed storage method based on source code semantic partitioning, and particularly relates to the technical field of cloud data distributed storage. The method comprises the steps of performing semantic partitioning on a source code, and segmenting the source code into a plurality of semantic blocks according to dimensions such as functional semantics, an abstract syntax tree structure, author information and version information; generating metadata containing information such as grammar type tags, file paths, line number ranges, author identifiers, version identifiers and access popularity for each semantic block; constructing a weighted directed acyclic graph (DAG) based on the semantic chunks and the dependency relationship thereof; superposing a metadata layer in the DAG structure, and recording information such as function call dependency, inter-block reference relationship and version evolution chain; blocks with relatively high access frequency and close semantics are aggregated into super blocks, the traversal depth is reduced, and meanwhile, hot data and cold data are differentiated for hierarchical storage by adopting a cold and hot data management strategy; and evaluating a parent block aggregation degree through a BDS algorithm, determining a block sorting priority, and optimizing super block boundary division. Compared with the prior art, the method has the advantages that the semantic retrieval efficiency, the incremental updating capability and the distributed query performance of the source code storage system are improved.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Game strategy retrieval method and device based on event-driven knowledge graph embedding

The invention provides a game strategy retrieval method and device based on event-driven knowledge graph embedding. The method comprises the following steps: performing structured analysis on strategy data, and updating entities and relationships; performing graph embedding calculation on entities and relationships in the knowledge graph to obtain vector representation among the entities, and constructing a searchable vector index; vectorizing a strategy query request of a user, and performing similarity retrieval on the strategy query request and the vector index to obtain a first candidate entity set similar to the query vector; performing relation reasoning by taking the first candidate entity set as a starting point in the knowledge graph to obtain a second candidate entity set in semantic association with the first candidate entity set, and combining the second candidate entity set with the first candidate entity set to form a target entity set; and performing correlation sorting on the target entity set to obtain a strategy result. According to the method, the strategy data can be automatically and structurally managed, and deep semantic retrieval is supported, so that the accuracy of game strategy content retrieval is improved, and the dynamic expansion capability is achieved.
Owner:QINGFENG (BEIJING) TECH CO LTD

Large model Text-to-SQL conditional clause improvement method and system based on fuzzy recall field value enhancement

The invention relates to the technical field of natural language processing and database query, and particularly discloses a large model Text-to-SQL conditional clause improvement method and device based on fuzzy recall field value enhancement, which are used for solving the problem of SQL query errors caused by inaccurate field values input by a user. According to the main scheme, the method comprises the following steps: extracting and recognizing date, field name and field value keywords in user input through a large model entity; 2) dynamically acquiring related field values from a database by combining a multi-stage recall mechanism of ElasticSearch (ES) fuzzy matching and vector semantic retrieval, and improving the recall accuracy through threshold screening; (3) organizing recall field values into structured contexts, injecting the structured contexts into a large model cue word template, and guiding the structured contexts to generate SQL condition clauses (such as WHERE clauses) containing accurate field values; and (4) changing a real-time synchronous database into an ES and vector database to ensure data consistency.
Owner:SICHUAN UNIV

Listed company operation risk early warning method based on multi-source auditing and text semantic fusion

The invention discloses a listed company operation risk early warning method based on multi-source auditing and text semantic fusion, and relates to the technical field of auditing, and the method comprises the steps: S1, crawling and converging multi-source heterogeneous data of listed company financial newspapers, auditing suggestions, supervision announcements, inquiry letters, news public opinions and market transactions; according to the method, unstructured texts are subjected to cleaning, blocking and semantic vectorization processing, each text segment is embedded into a high-dimensional semantic space, a vector index is established, a bottom-layer knowledge base of an RAG framework is formed, in the stage, it is ensured that the data structure is uniform, the source is traceable, standardized input is provided for subsequent semantic retrieval and modeling, and the reliability of the system is improved. S2, a query expression is constructed based on a target company, a time window and a risk topic, dense semantic retrieval and sparse BM25 retrieval methods are comprehensively used, a time decay and source credibility weighting mechanism is introduced, and the problems that a traditional method is single in data dimension and information is split are solved.
Owner:NANJING UNIV OF FINANCE & ECONOMICS

Structured data retrieval system and method based on semantic matching and hierarchical indexing

The invention discloses a structured data retrieval system and method based on semantic matching and hierarchical indexing, and the related retrieval system comprises a first construction module which is used for extracting slice data in a preset vector library and meta-information corresponding to the slice data, and constructing a text node object containing an id; the second construction module is used for traversing a text node object to obtain meta-information subjected to hierarchical structure processing, and constructing a nested index tree; the directory decomposition module is used for receiving an input text, performing decomposition based on a hierarchical structure and generating a corresponding query vector; the retrieval module is used for performing semantic retrieval and hierarchical retrieval on the text in sequence to obtain a retrieval result; the grouping and sorting module is used for grouping the retrieval results according to the hit hierarchy, sorting the retrieval results in each group according to a descending order, and combining all groups to obtain a final retrieval result list; and the data backtracking module is used for acquiring an original text field from the vector database according to the id corresponding to the retrieval result.
Owner:BIAOYIZHONG DIGITAL TECHNOLOGY (ZHEJIANG) CO LTD

Power grid power transformation engineering knowledge graph construction and retrieval method and system

The invention relates to the technical field of electric power engineering information processing, and discloses a power grid power transformation engineering knowledge graph construction and retrieval method and system, and the method comprises the steps: carrying out the dynamic adaptive partitioning of a power grid power transformation engineering related document, and obtaining semantic coherent and independent text blocks; extracting entities and relationships based on the text blocks, and complementing implicit entities and relationships through a multi-round refining mode; performing fusion and disambiguation on the extracted and complemented entities and relationships to form a unified knowledge graph; performing hierarchical clustering on the formed knowledge graph to generate a multi-granularity community structure and a corresponding community report; intention resolution and pre-judgment guidance are carried out aiming at fuzzy questions of the user, and a retrieval strategy is optimized; and executing multi-hop semantic retrieval based on the optimized retrieval strategy, recalling related knowledge and generating answers. According to the method, automation, precision and intelligentization of power grid power transformation engineering knowledge graph construction and full-link retrieval can be realized.
Owner:SOUTHWEST ELECTRIC POWER DESIGN INST OF CHINA POWER ENG CONSULTING GROUP CORP

Cache-Generated Frequently Asked Questions Page

In one embodiment, a method for a cache-generated frequently asked questions page includes converting a received query into a set of embeddings and performing a semantic searching operation based on information contained in the set of embeddings against a cache of question and answer pairs. The method further includes returning a stored answer to the received query responsive to a determination that a particular question and answer pair in the cache of question and answer pairs meets a threshold similarity level for the information contained in the set of embeddings, the stored answer derived from the particular question and answer pair and performing a large language model operation to generate an answer to the query responsive to a determination that no question and answer pair in the cache of question and answer pairs meets the threshold similarity level for the information contained in the set of embeddings.
Owner:CISCO TECHNOLOGY INC

Water conservancy design file retrieval system and method based on local lightweight large model

The invention discloses a water conservancy design archive retrieval system and method based on a local lightweight large model, and the method comprises the steps: S1, constructing a Python automatic preprocessing assembly line, extracting texts for PDF and Word multi-format archives, correcting metadata, and outputting standardized data; s2, constructing a full-text retrieval and semantic retrieval dual-mode cross-document retrieval service by relying on a Weavi ate local vector database and a lightweight text embedding model; s3, analyzing a user query intention through a local large model, synchronously triggering metadata accurate retrieval and content semantic retrieval, and generating a structured result; and S4, integrating the core module into a local area network Web platform, adopting Docker containerization deployment, and combining an RBAC permission model and JWT authentication to guarantee security. The system comprises a preprocessing module, a cross-document retrieval module, an intelligent agent module and a background management module, and collaboration is achieved through a standardized API. According to the method, the problem of archive fragmentation is solved, multi-mode retrieval breaks through keyword limitation, an intelligent agent reduces manual intervention, a localized architecture prevents secret-related leakage, background management adapts to an existing I T environment, and full-process intelligent archive service is provided for water conservancy design.
Owner:ZHONGSHAN WATER CONSERVANCY PROJECT SURVEY & CONSULT CO LTD

Code generation and evaluation method and system based on RAG and multilevel decision tree

The invention provides a code generation and evaluation method and system based on RAG and a multilevel decision tree, and the method comprises the steps: integrating project related design documents, and constructing a knowledge base capable of semantic retrieval through a vectorization technology; associating business demand description with related documents in the knowledge base based on an RAG technology, and performing demand semantic enhancement to generate a technology demand cue word; receiving the technical requirement cue word by adopting a large language model so as to generate a complete code conforming to business logic; constructing a four-level decision tree evaluation system, and sequentially executing code quality scanning, deployability verification, dynamic test verification and demand satisfaction verification through an evaluation assembly line to generate an evaluation result; and generating an optimization suggestion according to the evaluation result so as to trigger an iteration generation process when the code does not pass the verification, thereby solving the problems of disjunction between code generation and business requirements, low verification efficiency and insufficient iteration optimization.
Owner:SHANDING YUNKE INFORMATION TECHNOLOGY CO LTD

Household appliance knowledge question-answering method and system based on retrieval enhancement generation

The invention provides a household appliance knowledge question-answering method and system based on retrieval enhancement generation. The method comprises the following steps: acquiring household appliance field multi-modal data from a multi-format document library; extracting text information, table information and chart information in the multi-modal data; performing domain term injection processing on the extracted information, and constructing a packet domain enhancement index; receiving a natural language question input by a user; the natural language problem is analyzed through a query optimizer, and semantic retrieval and keyword retrieval are executed in parallel; carrying out fusion processing on the semantic retrieval result and the keyword retrieval result; selecting matched document fragments by adopting a relevancy sorting algorithm; inputting the matched document fragments into a large language model to generate candidate answers; verifying the compliance and traceability of the candidate answers through a credibility evaluation module; outputting a final answer with a reference source; and storing the high-frequency questions and the final answers into a cache library to solve the problems that the answer accuracy of a knowledge question-answering system is reduced and the response efficiency is limited.
Owner:SICHUAN HONGMEI INTELLIGENT TECH CO LTD

Systems and methods for implementing permission bypass for large language models

Systems and methods are provided for implementing permission bypass for LLM applications. One system includes an electronic processor that may be configured to receive, from a user device of a user, a user query pertaining to a topic. The electronic processor may also be configured to determine, responsive to a semantic search of a vector database, a plurality of electronic files related to the topic of the user query. The electronic processor may also be configured to determine, based on a permission level of the user, a first portion of the plurality of electronic files, where the first portion of the plurality of electronic files are accessible to the user under the permission level. The electronic processor may also be configured to generate, using a LLM, a response to the user query based on the first portion of the plurality of electronic files.
Owner:BOOST SUBSCRIBERCO LLC

Semantic search in high-dimensional spaces using euclidean distance and cluster-based optimization

Computer-implemented systems and methods implement semantic search in high-dimensional vector spaces, specifically tailored for use with large language models (LLMs). In particular, clustering is combined with Euclidean distance measurements to facilitate real-time vector searches. By implementing clustering, the invention reduces the computational complexity and costs associated with Euclidean distance calculations, which are typically more resource-intensive than other methods such as cosine similarity. This reduction is achieved by limiting the scope of distance calculations to within clusters, thereby avoiding the inefficiencies and diminished accuracy otherwise encountered by existing systems when using Euclidean distance in high-dimensional spaces. As a result, the invention retains the benefits of Euclidean distance, such as its superior granularity and precision in measuring semantic relevance, without succumbing to the usual drawbacks of high computational demands and poor scalability.
Owner:AICEBERG INC