Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

41531results about "Digital data information retrieval" patented technology

Intelligent question answering method based on collaboration between large language model and knowledge graph

Provided in the present application is an intelligent question answering method based on a collaboration between a large language model and a knowledge graph, relating to the technical fields of artificial intelligence and natural language processing, the method comprising: decomposing a complex question into a plurality of simple questions, and analyzing the degree of association between the simple questions and a basic function so as to form a multi-hop reasoning path; automatically extracting structured information from the simple questions on the basis of a multi-task learning framework of a large model, so as to construct a knowledge graph; and constructing a cumulative reasoning learning framework on the basis of a logic reasoning large model, and performing iterative verification on a process result formed by the knowledge graph on the basis of the multi-hop reasoning path, so as to correct the reasoning path until a correct answer is inferred.
Owner:INSPUR GENERSOFT CO LTD

Power plant operation and maintenance knowledge intelligent query method based on large language model and RAG technology

The invention discloses a power plant operation and maintenance knowledge intelligent query method based on a large language model and an RAG technology. The method comprises the following steps: constructing a power plant operation and maintenance knowledge vector library covering structured, semi-structured and unstructured data; receiving a natural language question of a user, inputting an improved instruction to align a preprocessor, and generating a question semantic vector and an intention tag; relevant knowledge fragments are retrieved and sorted through a semantic matching retriever in combination with the intention labels; constructing a large language model cue word structure based on the retrieval result and the original question, generating candidate answers and recording a reference path; and finally, performing term specification and consistency verification according to the expert rule base, and outputting a structured and traceable final answer. According to the invention, the improved RAG technology is fused to realize intelligent query of the operation and maintenance knowledge of the power plant.
Owner:JIANGSU GUOHUACHENJIAGANG POWER GENERATION CO LTD

System and method for estimating confidence and implementing metacognitive abilities in artificial intelligence systems

In a described embodiment, a system for information processing is provided including a data acquisition module configured to receive feedback corresponding to one or more outputs generated by a language model. The system further includes a cognitive reasoning module configured to evaluate the reasoning process of the language model, emulate cognitive functions including metacognitive processes, and generate an assessment based on an analysis of the received feedback, wherein the assessment includes classifying the one or more outputs into components, assigning quality scores for each component, and identifying an improvement corresponding to the one or more outputs. Additionally, the system includes a process adjustment module coupled to the cognitive reasoning module for adjusting the reasoning process of the language model based on the assessment is provided. A refinement module coupled to the process adjustment module is provided for iteratively refining the reasoning process based on subsequent updates to the generated assessment until a performance threshold is met.
Owner:BLACKBERRY LTD

Method and system for realizing Text2SQL (Structured Query Language)

The invention discloses a Text2SQL (Structured Query Language) implementation method and system, and relates to the field of data processing, and the method comprises the following steps: firstly, receiving a natural language query, and analyzing a query intention, field classification and a key entity through a planner; the searcher obtains domain knowledge, entity information, a database table structure and a historical query mode in a multi-path parallel mode based on the planning result; the generator constructs an SQL framework according to the retrieval result and generates an initial statement; the verifier carries out grammar, table field, authority and logic multi-dimensional verification on the SQL, and if the verification fails, iteration adjustment is carried out to generate logic; when the SQL is executed, the result is formatted and a natural language explanation containing query logic, a data source and a calculation method is generated if the SQL is executed successfully, and a diagnosis and error correction mechanism is started for correction and then rechecking is performed if the SQL is executed unsuccessfully. According to the method, through deep fusion of domain knowledge, whole-process verification error correction and interpretability enhancement, the accuracy, robustness and user interaction experience of SQL conversion in a professional scene are improved.
Owner:XUNTU TECH (SHANGHAI) CO LTD

Multi-element sales planning agent system and method

The invention discloses a multi-element sales planning agent system and method, and aims to improve the intelligence and precision of sales planning. The system comprises a collection module, an analysis module, an optimization module, a creation module and a generation module. The collection module is used for receiving multi-modal data such as marketing targets and extracting key marketing elements. The analysis module is used for generating a target user portrait and extracting marketing strategy analysis data. And the optimization module is used for calculating a medium putting weight by utilizing reinforcement learning and generating a medium strategy scheme. And the creation module generates a propagation theme and marketing content by adopting a generative artificial intelligence technology. And the generation module predicts a delivery effect by using a machine learning model and dynamically optimizes a medium strategy and a content scheme. Through multi-modal data fusion, intelligent analysis and optimization, closed-loop processing from data acquisition to marketing execution is realized, the marketing decision-making efficiency is improved, and brand promotion accuracy and market adaptability are enhanced.
Owner:SUZHOU DUOYUAN DATA CO LTD

Cross-modal retrieval method for semantic and vector fusion in data space

The invention provides a cross-modal retrieval method for semantic and vector fusion in a data space, which belongs to the field of cross-modal information retrieval, and comprises the following steps: firstly, collecting and preprocessing multi-modal data; generating modal embedding and storing by utilizing the pre-training model; a shared semantic space is constructed, cross-modal vector alignment is optimized through comparative learning, and a modal mapping network is designed to enhance the embedding projection effect; storing the aligned embedding by using a Milvus database, and constructing an HNSW index; user text or image query is processed, text query analyzes limiting conditions to generate enhanced embedding, and image query extracts characters through OCR and fuses the characters with image features to generate embedding; in a database, through condition screening and semantic similarity calculation, a Top-K candidate item is retrieved; performing multi-modal correlation sorting on the candidate results and returning the results; according to the method, the shared semantic space is constructed, the alignment effect of different modal embedding is optimized, efficient storage and index management of multi-modal embedding are carried out, and real-time retrieval of large-scale cross-modal data is achieved.
Owner:HARBIN ENG UNIV

Explanatory model architecture for image scoring reasoning

A method includes obtaining an image, the image associated with a mask corresponding to a portion of the image, generating a plurality of images based on the image and the mask, each image of the plurality of images depicting a different color in the portion of the image corresponding to the mask, executing a machine learning model to generate an image performance score for each of the plurality of images, ranking the plurality of images according to the image performance scores for the plurality of images, and generating a record comprising one or more images of the plurality of images based on the rankings of the plurality of images.
Owner:VIZIT LABS INC

Advanced systems and methods for multimodal ai: generative multimodal large language and deep learning models with applications across diverse domains

Systems and methods are provided for improving generative artificial intelligence (AI). Systems and methods can integrate more reliable data sources and enhance generative AI training and inference processes for complex tasks. The integration of real-time data and expert input can be included as crucial steps in aligning AI outputs with improved accuracy. Similarly, fine-tuning methodologies and augmentation algorithms can be used to focus on minimizing the occurrence of fabricated content, thereby significantly increasing the chances that the information generated is both current and credible.
Owner:UNIV OF MIAMI

NL2SQL optimization method and device based on large model, equipment and medium

The invention discloses an NL2SQL optimization method and device based on a large model, equipment and a medium, and relates to the technical field of artificial intelligence, the method comprises the following steps: constructing a target metadata knowledge base, and obtaining an initial natural language query request; determining each target entity corresponding to the initial natural language query request, and determining missing target SQL elements in the initial natural language query request based on each target entity; generating a first cue word based on the initial natural language query request, the target SQL element and the target metadata knowledge base, and complementing the target SQL element based on the first cue word by utilizing the target large model to obtain a target natural language query request; and generating a plurality of candidate SQL statements corresponding to the target natural language query request by using the target large model, verifying each candidate SQL statement, and determining a target SQL statement from each candidate SQL statement based on a verification result. According to the method, the accuracy of the NL2SQL can be improved by utilizing a large model.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Efficient remote pointer sharing for enhanced access to key-value stores

A method to share remote DMA (RDMA) pointers to a key-value store among a plurality of clients. The method allocates a shared memory and accesses the key-value store with a key from a client and receives an information from the key-value store. The method further generates a RDMA pointer from the information, maps the key to a location in the shared memory, and generates a RDMA pointer record at the location. The method further stores the RDMA pointer and the key in the RDMA pointer record and shares the RDMA pointer record among the plurality of clients.
Owner:IBM CORP

Cross-modal knowledge reasoning method based on multi-modal large model

The invention relates to a cross-modal knowledge reasoning method based on a multi-modal large model. In a cross-modal knowledge reasoning process, an existing model is usually limited by single-modal information extraction and shallow feature fusion, so that deep semantic association among data such as texts, images and videos is difficult to fully capture. In order to solve the problem, the invention provides a model for fusing multi-modal information such as texts, images, videos, documents and the like, and processing of multi-modal data is converted into unified feature extraction, interaction and deep reasoning tasks by fully utilizing a supervision fine tuning strategy, a self-adaptive attention mechanism and a cross-language processing technology. The model adopts a modular design, integrates multi-source data complementary analysis, spatial-temporal feature modeling and emotional semantic analysis, and realizes multi-modal collaborative interaction, dynamic scene understanding, long video key event analysis and man-machine co-emotional response. Through sufficient training, the multi-modal large model shows excellent logical reasoning ability and emotion understanding ability in a complex cognitive task, and a brand new solution is provided for efficient extraction, deep semantic analysis and intelligent response of cross-modal information.
Owner:SHENYANG INST OF COMPUTING TECH CO LTD THE CHINESE ACAD OF SCI

Intelligent tracing method and system for agricultural non-point source pollution based on knowledge graph

The invention relates to the technical field of agricultural traceability, and discloses an agricultural non-point source pollution intelligent traceability method and system based on a knowledge graph, and the method comprises the steps: achieving the system integration of pollution source features through constructing the knowledge graph fusing multi-dimensional data, and building a sky-ground integrated monitoring network to obtain multi-scale dynamic data. An intelligent traceability mechanism is formed based on deep coupling of a knowledge graph and monitoring data, a high-precision pollution identification model is trained in combination with historical data, and real-time response and dynamic traceability of an over-standard pollution area are achieved. And verifying a traceability result through feature matching and semantic reasoning, quantitatively calculating a pollution contribution rate, and finally generating a reliable traceability conclusion containing the position, the type and the contribution rate. According to the scheme, the bottlenecks of data fragmentation, monitoring simplification, extensive analysis and the like of a traditional method are broken through, and the accuracy, timeliness and credibility of pollution traceability are remarkably improved through a series of full-chain technical systems of traceability identification, analysis and positioning.
Owner:NANJING ACAD OF ENVIRONMENTAL PROTECTION SCI

Intention classification method and device based on vector retrieval and context awareness and medium

The invention discloses an intention classification method and device based on vector retrieval and context awareness and a medium, and relates to the technical field of artificial intelligence. The method comprises the steps of extracting business metadata and associating the business metadata with typical problem examples to generate a standardized service description document; encoding the standardized service description document into a high-dimensional semantic vector through a pre-training language model, and constructing a neighbor search index to store the high-dimensional semantic vector; splicing the user identity information and the current question text into an enhanced query statement, and encoding the enhanced query statement into a context-aware dynamic query vector through a semantic model; performing similarity retrieval based on the dynamic query vector to obtain candidate intelligent services, performing business domain filtering, context weighted sorting and dynamic priority rearrangement, and outputting target recommendation services; by collecting interactive behavior data of a target recommendation service, quality scoring is performed on service descriptions and problem examples based on a preset evaluation rule, and the service descriptions and the problem examples of which the quality scores are lower than a quality threshold value are updated.
Owner:INSPUR GENERSOFT CO LTD

Building electromechanical BIM model information rapid retrieval method and system

The invention discloses a building electromechanical BIM model information rapid retrieval method and system, and the method comprises the steps: generating composite retrieval parameters fusing semantic keywords and three-dimensional coordinate constraints according to a multi-mode retrieval instruction inputted by a user; on the basis of the composite retrieval parameters, constructing a dynamic search space by utilizing a hierarchical graph convolutional network, and generating a candidate model index structure of multi-dimensional feature coding; inputting the candidate model index into a multi-objective optimization engine, performing real-time optimization on a search path by adopting a dynamic pruning algorithm driven by reinforcement learning, and outputting a candidate model set of which the confidence coefficient is higher than a preset confidence threshold after pruning; and on the basis of the candidate model set, associated equipment nodes are expanded through a knowledge graph embedding and complementing technology, and an enhanced retrieval result set containing the hidden associated equipment is generated. By utilizing the embodiment of the invention, efficient, multi-dimensional and multi-modal information accurate positioning and quick retrieval can be realized in a large-scale complex BIM model.
Owner:杭州美屋美居数智科技有限公司

Automated Prompt Augmentation And Engineering Using ML Automation In SQL Query Engine

A database system generates a prompt for an LLM or other machine learning (ML) model to narrow the search space to highly relevant information about a database. A distinct instance of a classifier, a clustering algorithm, or a topic modeling model can be trained based on information from ML automation within the database system, respectively for each column or table in the database. Model instances can then be used during generative LLM inferencing to identify relevant sources of data to answer the user's question. Thus, the prompt generation combines ML automation and other ML models or an LLM for topic modeling and schema description.
Owner:ORACLE INT CORP

Digital integrated quality management system based on multi-source data fusion

The invention relates to a digital integrated quality management system based on multi-source data fusion, and belongs to the technical field of industrial internet and quality management. A data acquisition layer of the system obtains real-time and static multi-source heterogeneous data through a multi-source adapter; the data processing layer is used for cleaning, converting and standardizing the acquired data; the intelligent analysis layer performs deep analysis and prediction on the data by using an adaptive quality prediction model, an anomaly detection module and a root cause analysis engine; the application service layer displays a quality trend and an anomaly detection result through a visual billboard, and provides credible tracing and collaborative decision-making functions; and the feedback closed layer adjusts system processing logic according to the decision support data to form closed-loop quality control. According to the method, real-time fusion and efficient utilization of multi-source data are realized through a dynamic routing technology, an adaptive quality prediction model and a block chain evidence storage mechanism, and the intelligent level and decision-making efficiency of quality management are remarkably improved.
Owner:CHONGQING BOJUN IND TECH CO LTD

SQL intelligent generation method and system for business query

The invention provides an intelligent SQL generation method and system for business query, and belongs to the technical field of artificial intelligence. Related data of query statements are acquired, and a corresponding query intention knowledge graph is constructed by semantic clustering; the method comprises the following steps: analyzing a historical SQL statement structure, extracting a natural language template and an SQL template, and expanding through a large language model to generate a feed-shot example set; and constructing a composite cue word template by combining task setting guidance, a feed-shot example and CoT chain thinking reasoning guidance. An intention completion module is arranged in a large language model, a natural language query statement of a user is combined with a composite cue word template context, a structured query statement is generated through entity recognition, semantic completion, parameter filling and fuzzy intention training, and the structured query statement is converted into a standard SQL statement through a knowledge graph and a template. According to the method, the use threshold of business personnel is remarkably reduced, and efficient conversion from natural language questions to SQL statements is realized.
Owner:国网福建省电力有限公司营销服务中心 +1

Enhancing retrieval augmented generation accuracy

Provided is a method including obtaining a prompt, determining a prompt embedding vector representing the prompt in an embedding space, modifying the prompt embedding vector using a trained model configured to adjust prompt embedding vectors to decrease proximity to vectors of blocks in a data set from which data is retrieved to augment generation by the generative AI model, determining that the modified prompt embedding vector is within a threshold distance to vectors in the embedding space corresponding to one or more blocks in the data set, selecting the one or more blocks in the data set, generating a response using the generative AI model based on the selected one or more blocks in the data set, quantifying an amount of influence of the respective block on corresponding text in the generated response, and providing the response and a representation of the quantified amount of influence as an output.
Owner:TELPERIAN INC

Modular ai agent system with dynamic skill registry and resource management for enterprise applications

Systems and methods for integrating generative artificial intelligence (AI) within Software-as-a-Service (SaaS) platforms to automate data operations, synchronize cross-platform workflows, and enable intent-based interactions. A platform displays table structures of items and characteristics linked to a common objective, provides input interfaces, and enrolls AI agents as credentialed users with read / write privileges. The system prompts agents with column types, structural relations, and role profiles to generate and execute editing instructions that progress workflow objectives, detect missing or inconsistent data, and notify users or request information as needed. Hierarchical access schemes permit multiple agent instances with inherited privileges and resource limits managed through an AI center. Agents can operate as autonomous team members, analyze outputs, and support natural-language explanation sessions. Additional embodiments coordinate inter-service updates, maintain deviation detection tools, and construct tailored products and platform elements. These capabilities improve robust automation, decision support, and operational efficiency in complex SaaS environments.
Owner:MONDAY COM LTD

Intelligent sound box voice processing method and system based on artificial intelligence

The invention provides an intelligent sound box voice processing method and system based on artificial intelligence, and the method comprises the steps: obtaining audio data and mouth shape video data, carrying out the processing of the audio data and the mouth shape video data, and carrying out the multi-modal feature fusion, and obtaining a fusion feature; performing bimodal voice activity detection on the fusion features to obtain effective voice data; performing context sensing recognition of audio and video fusion on the effective voice data to obtain a first text; constructing a user feature model, and performing semantic understanding on the text based on the model to obtain an understanding result; performing intention recognition and slot filling based on the understanding result to obtain user intention and key information; generating a response strategy in combination with the user intention, the key information and the environment perception data; generating response voice according to the response strategy; and monitoring feedback information of the user to the response voice in real time, and updating the user feature model and the response strategy evaluation model based on feedback. According to the scheme, the voice can be recognized more accurately, and the safety and robustness of the system are enhanced.
Owner:SHENZHEN ZHANDIAN SMART TECH CO LTD

Frequency converter fault prediction method and system based on machine learning

The invention relates to the field of frequency converter fault detection, and discloses a frequency converter fault prediction method and system based on machine learning, and the method comprises the steps: obtaining multi-dimensional real-time data in the operation process of a frequency converter; constructing a dynamic mapping relation to obtain a basic feature set; generating a time sequence feature vector capable of reflecting the state change of the equipment based on the basic feature set; comparing, analyzing and judging whether the equipment state deviates from a normal operation interval or not based on the historical operation data and the time sequence feature vector, and outputting a state deviation index; performing abnormal fluctuation judgment on the time sequence feature vector; extracting fluctuation amplitude and frequency characteristics of the key indexes to obtain quantitative description data of abnormal fluctuation; inputting the quantitative description data of the abnormal fluctuation into an abnormal prediction model; and generating a coping strategy and a triggering condition of the coping strategy based on the risk prediction result. The method has the advantages that the abnormal state of the frequency converter is recognized in time, and potential risks are predicted.
Owner:SHENZHEN ZHONGDA ELECTRIC TECH CO LTD

Multi-stage LLM with unlimited context

A system and method for efficient natural language processing combines large and small language models with a thought caching architecture. The system includes a router that directs prompts either to a large language model for thought generation or to a thought cache containing previously generated thoughts. When using the large model, generated thoughts are combined with the original prompt and routed through a smaller language model to produce responses. The thought cache stores reasoning patterns that can be retrieved and reused, eliminating the need to regenerate similar thoughts for related prompts. The system supports both local and cloud-based caching, enabling personal and enterprise-wide thought storage and retrieval. This architecture reduces computational overhead while maintaining reasoning capabilities, effectively extends context windows beyond traditional limits, and enables efficient scaling across different deployment scenarios. The system can operate with reduced resources by leveraging cached thoughts without requiring constant access to the large model.
Owner:ATOMBEAM TECH INC

RAG enhanced Text-to-SQL query method and system for large-scale database environment

The invention discloses an RAG enhanced Text-to-SQL (Structured Query Language) query method and system for a large-scale database environment. The method comprises the following steps: constructing a vector database; based on the original query of the user, generating query enhancement description aiming at Schema recall and query enhancement description aiming at SQL (Structured Query Language) by utilizing LLM (Logistics Language Model); carrying out Schema recall and historical question and answer pair recall operations in a double-way parallel manner by utilizing an RAG technology and relying on the constructed vector database; based on query enhancement description, recalled Schema and historical question and answer pairs, generating an SQL query statement by using LLM in combination with RAG, and realizing Text-to-SQL conversion; the generated SQL query statement is subjected to post-processing optimization through a database interface, the post-processing optimization comprises grammar verification, performance optimization and error correction, the SQL query statement is combined for execution and feedback, efficient, accurate and reliable natural language query is achieved, and continuous optimization can be conducted through user feedback.
Owner:COSCO SHIPPING TECH CO LTD

Test analysis and report generation method and system based on large model retrieval enhancement

The invention relates to the technical field of automatic report generation, and provides a test analysis and report generation method and system based on large model retrieval enhancement, and the method comprises the steps: firstly carrying out the semantic task decomposition of user query, and meanwhile, achieving the multi-dimensional information extraction through the butt joint of a knowledge vector library and a structured image-text knowledge system established by a private knowledge base module. And three strategies of RAG, Self-RAG and Graph-RAG are combined to enhance the generation capability of the large model so as to accurately obtain background knowledge. And constructing a standard SQL statement according to a structured query requirement, and performing data filling according to template prompt by an output result fusion module in combination with a query background and an SQL execution result to form a test report. By optimizing a retrieval enhancement generation method, professional knowledge supplement related to query is realized, and the reasoning ability of a large model in a professional scene is effectively enhanced, so that the accuracy and the intelligent level of data analysis are improved.
Owner:CHINA ELECTRONIS TECH INSTR CO LTD

Resource recommendation method and system based on hybrid retrieval RAG

The invention relates to the technical field of intelligent recommendation, and discloses a hybrid retrieval RAG-based resource recommendation method and system, and the method comprises the steps: collecting resource text data, and constructing a vector library and a tag library; expanding the user question based on the language model to obtain a plurality of semantic extension questions; performing intention recognition, judging whether the user question is a resource recommendation question, and if yes, determining a target classification type; screening the data according to the field definition in the tag library to obtain a candidate knowledge fragment set; obtaining candidate vectors, mapping the user question and the semantic extension question into query vectors, calculating the similarity between the query vectors and each candidate vector, and selecting knowledge supplement content; and performing resource splicing on all the knowledge supplement contents to generate resource recommendation answers. According to the method, a structured label screening mechanism and a semantic vector fine arrangement mechanism are fused, and the problems of recall redundancy, matching deviation and the like caused by the fact that an existing RAG system only depends on semantic similarity retrieval are solved.
Owner:ZHEJIANG DAGU TECH CO LTD

Security Methods and Systems for Multi-Agent Generative AI Applications

A system and method for securing inter-process communications (IPCs) between generative AI and external tool servers, including intercepting IPCs having a requested operation, performing a security analysis on the IPCs for security threats, performing a permission validation for permissions for the requested operation of each IPC, and either approving or blocking the requested operation of each IPC based on the security analysis and the permission validation.
Owner:MADISETTI VIJAY

Method and system for dynamically calling Java statistical analysis interface based on LLM and MCP

The invention discloses a method and system for dynamically calling a Java statistical analysis interface based on LLM and MCP protocols, belongs to the technical field of data processing and analysis technology, artificial intelligence technology and application program interface (API) integration, and aims to solve the technical problem of how to improve the response speed and flexibility of statistical analysis requirements. In order to reduce the development and maintenance cost of a statistical analysis function, the adopted technical scheme is as follows: a Java annotation mechanism is utilized to mark and describe an existing Java statistical analysis system, a Java method of which the annotation is defined is obtained, and the Java method of which the annotation is defined can be automatically identified by an MCP server and registered as an MCP tool; the MCP server follows an MCP protocol and communicates with the LLM serving as an MCP client, and the LLM dynamically discovers available MCP tools through the MCP protocol; after natural language query of a user is understood through LLM, one or more appropriate MCP tools are intelligently selected, calling parameters are generated, and an MCP server is requested to execute through an MCP protocol.
Owner:SHANDONG INSPUR E-GOVERNMENT SOFTWARE LTD

Using Machine Learning Techniques To Improve The Quality And Performance Of Generative AI Applications

PendingUS20250284721A1Digital data information retrievalCommerceDatabase machineObject store
A database system integrates in-database machine learning (ML) models with in-database large language models (LLMs) or other generative artificial intelligence (AI) models that enable new applications. The database system receives one or more inferences from an ML model and provides an inference input to a retrieval agent of an object store. One or more vector stores represent a plurality of reference documents using semantic encodings. The retrieval agent performs a similarity search of the one or more vector stores to retrieve a set of passages from the plurality of reference documents based on similarity of encodings of the inference input and encodings of passages in the plurality of reference documents. The database system generates a linguistic prompt for an LLM having a context including the inferences and passages and applies the LLM to the linguistic prompt to generate a natural language explanation of the one or more inferences.
Owner:ORACLE INT CORP

SQL generation method and apparatus based on large language model, device and storage medium

A SQL generation method based on a large language model, comprising: generating data definition language (DDL) prompt slot information and data example slot information on the basis of a configuration operation of a user on a SQL database; receiving a natural language query request input by the user, and transcribing the natural language query request to obtain question transcribing slot information; on the basis of the question transcribing slot information, the DDL prompt slot information and the data example slot information, acquiring complete prompt information matched with a large language model; and inputting the complete prompt information into the large language model to generate an executable SQL statement matched with the natural language query request.
Owner:DATAGRAND TECH INC

Question and answer reasoning method and device based on key value cache compression, equipment and medium

The invention discloses a question and answer reasoning method and device based on key value cache compression, equipment and a medium, and relates to the technical field of natural language processing, and the method comprises the steps: segmenting a cue word in a current question and answer task into a lexical sequence, and generating an initial key value cache of the lexical sequence; dividing the lexical element sequence into context lexical elements and tail end lexical elements corresponding to each layer based on a preset tail end window size of each attention layer of the target large language model; screening out keyword elements of each attention layer from the context lexical elements according to importance scores between key matrixes of the context lexical elements and query matrix mean values of the tail end lexical elements; removing key value pairs of lexical elements except the keyword elements in the initial key value cache to obtain a compressed key value cache; and generating a reasoning result corresponding to the compressed key value cache by using the target large language model. The high computing power consumption of the large language model caused by key value cache data increase is reduced, and the dependence of an existing key value cache compression method on a complete attention weight matrix is broken through.
Owner:ZHEJIANG TONGHUASHUN INTELLIGENT TECH CO LTD