Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1686 results about "List" patented technology

In computer science, a list or sequence is an abstract data type that represents a countable number of ordered values, where the same value may occur more than once. An instance of a list is a computer representation of the mathematical concept of a finite sequence; the (potentially) infinite analog of a list is a stream. Lists are a basic example of containers, as they contain other values. If the same value occurs multiple times, each occurrence is considered a distinct item.

Intelligent document analysis method and system

The invention discloses an intelligent document analysis method and system, the method is executed by the intelligent document analysis system, and the method comprises the following steps: carrying out layout analysis on a PDF page by adopting a deep learning model; merging the block list from bottom to top by adopting a recursive algorithm; carrying out balance optimization on the binary tree structure; and outputting a result of the processed binary tree structure by adopting a preorder traversal mode. Layout analysis is carried out by adopting a deep learning model, various complex typesetting formats such as multi-column layout, image-text mixed typesetting, tables, lists and the like can be effectively identified and processed, and semantically continuous text blocks are ensured to keep continuity in a tree structure through a tree structure optimization module; a global semantic error correction module is added to carry out global document semantic representation learning and carry out adaptive adjustment and error correction on a preliminary structure, so that deep ambiguity is eliminated, logic errors are repaired, and the consistency of a final analysis result and human reading logic on the semantic level is maximized.
Owner:SHANGHAI YILIAN INTELLIGENT TECH CO LTD

Data platform metadata automatic generation method based on large model

The invention relates to the technical field of metadata generation, and discloses a data platform metadata automatic generation method based on a large model, which comprises the following steps of: firstly, performing unified standardization and preliminary grammatical analysis on an original SQL (Structured Query Language) code to obtain a structured intermediate representation; on the basis, preliminary blood relationship analysis based on rules is carried out, and simple column references are quickly identified and processed. For complex expressions which are difficult to accurately analyze by a traditional method, code snippets and context information of the complex expressions are accurately extracted and submitted to a large language model for deep semantic understanding and complex blood relationship analysis. And finally, integrating the complex consanguinity analyzed by the large model with the initial consanguinity list to form a comprehensive and accurate field-level consanguinity, and further generating complete data platform metadata. In this way, the defect that a traditional analysis tool understands complex semantics is effectively overcome, and the accuracy and integrity of metadata generation are remarkably improved.
Owner:ZHEJIANG NON-LINEAR DIGITAL TECH CO LTD

Root cause analysis method and device based on large language model, equipment and medium

The invention discloses a root cause analysis method and device based on a large language model, equipment and a medium. The method comprises the steps that a target entity and a topological relation are extracted in response to a root cause analysis request; searching matched historical root cause cases in a vector database, inputting the topological relation, the historical root cause cases and cue words into a large language model to generate a doubtful point list, and extracting downstream entities in the doubtful point list; calling a detection tool to collect entity diagnosis data and identify abnormal downstream entities; returning to execute the operation of retrieving the historical root cause case until a preset iteration ending condition is met; and inputting the final doubtful point list, the abnormal diagnosis data, the topological relation and the historical root cause case into the large language model again to obtain a root cause reasoning result and generate a root cause report. According to the embodiment of the invention, through the full-chain design of natural language understanding, topological constraint, historical cases, dynamic detection and iterative reasoning, the fault positioning efficiency and the root cause accuracy are improved, and the operation and maintenance labor cost and the service fault time consumption are remarkably reduced.
Owner:BEIJING YOUTEJIE INFORMATION TECH

Multi-dimensional natural resource intelligent monitoring method and system based on big data analysis

The invention discloses a multi-dimensional natural resource intelligent monitoring method and system based on big data analysis, and relates to the technical field of resource monitoring, and the method comprises the steps: calculating the information sharing degree between different data streams, constructing an abnormal feature credibility distribution map, and carrying out the calculation of the abnormal feature credibility distribution map; and inputting the abnormal feature credibility distribution graph and the multi-source auxiliary information into an integrated classifier, carrying out joint analysis and comprehensive research and judgment on multi-dimensional information to output a target list, constructing an event response relation graph reflecting a monitoring task priority and an execution logic relation according to the attribute features of each element in the target list, and carrying out event response analysis on the monitoring task priority and the execution logic relation. Wherein each node represents a to-be-monitored target, an edge represents a dependency relationship between tasks, and an optimal monitoring execution path is planned for various intelligent monitoring carriers by using an improved A algorithm integrated with a multi-factor cost function. Through natural resource monitoring of multi-modal data interference correction, intelligent classification and path optimization and dynamic response, the monitoring quality and the execution efficiency are effectively improved.
Owner:JIANGXI GANDIYUAN TECHNOLOGY CO LTD

Information retrieval system and method based on semantic normalization

The invention discloses an information retrieval system and method based on semantic normalization, and relates to the technical field of artificial intelligence information, and the method comprises the steps: collecting a semantic query record input by a user, carrying out the preliminary semantic analysis, and generating structured data; on the basis of the structured data, entity disambiguation is carried out by utilizing a knowledge graph, abstract classes are generated through a neural network, calibration and dynamic weight adjustment are carried out, and high-confidence entity abstract classes and confidence scores are generated; entity abstract classes and confidence scores are combined with user contexts, an action-value function is calculated through a value network, and an optimal action is selected by utilizing a-greedy algorithm; executing semantic normalization mapping according to the optimal action, and obtaining an intermediate expression by using a meta-symbol dynamic generator; and performing index retrieval and multi-dimensional sorting based on the intermediate expression to generate a sorted retrieval result list. According to the method, the semantic fragmentation problem of multi-modal query is solved, and deep semantic alignment and dynamic weight calibration of heterogeneous data are realized.
Owner:上海笑聘网络科技有限公司

Structured data retrieval system and method based on semantic matching and hierarchical indexing

The invention discloses a structured data retrieval system and method based on semantic matching and hierarchical indexing, and the related retrieval system comprises a first construction module which is used for extracting slice data in a preset vector library and meta-information corresponding to the slice data, and constructing a text node object containing an id; the second construction module is used for traversing a text node object to obtain meta-information subjected to hierarchical structure processing, and constructing a nested index tree; the directory decomposition module is used for receiving an input text, performing decomposition based on a hierarchical structure and generating a corresponding query vector; the retrieval module is used for performing semantic retrieval and hierarchical retrieval on the text in sequence to obtain a retrieval result; the grouping and sorting module is used for grouping the retrieval results according to the hit hierarchy, sorting the retrieval results in each group according to a descending order, and combining all groups to obtain a final retrieval result list; and the data backtracking module is used for acquiring an original text field from the vector database according to the id corresponding to the retrieval result.
Owner:BIAOYIZHONG DIGITAL TECHNOLOGY (ZHEJIANG) CO LTD

Localized bidding document error checking method based on large language model

The invention relates to the technical field of text inspection and analysis, in particular to a localized bidding document error inspection method based on a large language model, which comprises the following steps of: constructing a bidding requirement knowledge graph which comprises a plurality of requirement item nodes, the requirement item node attribute comprises requirement content, a chapter to which the requirement item node attribute belongs and an importance level; establishing a preliminary mapping relationship between the requirement item node and the text content set, and generating a structured bidding file parse body; and carrying out double-round model analysis and error checking, generating an error record for the items judged to be not satisfied, and outputting a structured error list comprising the non-satisfied items and reasons of the non-satisfied items. The method improves the accuracy and coverage of error checking, and is especially suitable for checking scenes of bidding and tendering documents with numerous and jumbled contents and various formats.
Owner:SHANGHAI BELDEN PROJECT MANAGEMENT CONSULTING CO LTD

File uploading attack interception method based on semantic entropy enhancement

The invention provides a file uploading attack interception method based on semantic entropy enhancement, and aims at overcoming the defects of an existing file uploading security protection technology in the face of complex attacks. The method specifically comprises the steps that S1, file format analysis and content extraction are conducted, hidden scripts are mined through nested content recognition, and intermediate representation is generated through grammar cleaning and coding specifications; s2, constructing an abstract syntax tree and semantic entropy calculation, tracking a pollution chain, analyzing a high-risk function, identifying a high-entropy character string, modeling and controlling flow complexity, and generating a semantic entropy vector; s3, dynamic scoring and decision making are carried out, and accurate judgment is carried out in combination with white list perception, feature comparison, multi-modal model scoring, adaptive threshold and sandbox observation; and S4, carrying out real-time interception and feature synchronization, blocking malicious file landing, generating an attack log and synchronizing an attack fingerprint. The method takes the semantic entropy vector as a core, breaks through the limitation of static features, remarkably improves the recognition rate of complex attacks, reduces missed judgment, and guarantees the safety of Web applications.
Owner:CHINA LIFE INSURANCE CO LTD

Mixed retrieval method and system for multi-dimensional heterogeneous knowledge recall enhancement

The invention relates to the technical field of information retrieval, and provides a multi-dimensional heterogeneous knowledge recall enhanced hybrid retrieval method and system.The method comprises the steps that texts recalled through keyword retrieval, sparse vector retrieval and dense vector retrieval are screened through a reciprocal sorting fusion algorithm, and a text type candidate knowledge list is obtained; based on user query, generating and checking a query statement through a large language model, and retrieving an entity-relationship-attribute triple from the knowledge graph library; based on user query, generating enhanced knowledge through a knowledge graph enhanced retrieval method fusing keyword retrieval, vector retrieval and community retrieval; and carrying out format alignment and duplicate removal on the text type candidate knowledge list, the triple result and the enhanced knowledge to form a multi-modal candidate pool, and carrying out reordering score calculation and ordering on each piece of recall knowledge in the multi-modal candidate pool through a reordering model and a business rule to obtain a final retrieval result. And the coverage blind area of single retrieval on heterogeneous knowledge is solved.
Owner:DAREWAY SOFTWARE

Intelligent knowledge base management system and method based on large model

The invention belongs to the technical field of database management, and particularly relates to a knowledge base intelligent management system and method based on a large model, and the system collects question and answer texts and index information of different target subjects through a data collection module, and generates a first text hierarchical sequence and a directed index graph through word segmentation preprocessing and directed graph neural network mapping; the feature extraction module extracts text and index embedding representation in combination with a feature distillation model, and obtains similar overlapping degrees of different target main bodies through a three-dimensional overlapping degree model; the conflict simulation module constructs an initial knowledge base based on a tree database, detects index conflicts in target subjects and among the subjects by using a Bayesian simulation algorithm, and generates a conflict risk list; the conflict resolution module outputs a resolution strategy in combination with a conflict resolution confrontation model, and the feedback adjustment module iteratively optimizes the knowledge base until a conflict threshold and a user demand are met; according to the method, intelligent management of cross-subject question and answer indexes is realized, and knowledge base conflict resolution efficiency and cross-subject query accuracy are improved.
Owner:KAIENTAI (NANJING) TECH CO LTD

Large model-based standardized data processing method, electronic equipment, storage medium and computer program product

The invention provides a standardized data processing method based on a large model, electronic equipment, a storage medium and a computer program product, and the method comprises the steps: generating a standard field semantic vector through large model coding field name character string input, a field value sample list and context structure information of a field; calculating a semantic offset degree between the semantic vector of the field to be standardized and the semantic vector of the standardized field, and generating a successfully matched field pair set and a normalized failed field set; obtaining a historical standard version, and generating an optimal field and maximum score version mapping set; semantic positioning is carried out in the maximum score version according to the optimal field, and a matching field is calculated; and obtaining standard mapping field information under the current version based on the mapping rule table of the matching field, and outputting the standard mapping field information. The problem that a traditional rule system depends on a fixed field and is not sensitive to version change is solved.
Owner:CHINA NAT INST OF STANDARDIZATION

Script preservation method, device, equipment, medium and program product

The invention provides a script preservation method which can be applied to the technical field of artificial intelligence. The script preservation method comprises the following steps: carrying out semantic understanding on demand related files by utilizing a large model, extracting core elements, and generating a structured demand change list; analyzing to obtain an abstract syntax tree of an existing script code, and constructing a mapping relationship between the software automation test element and a script code block; dynamically comparing the structured demand change list with an existing script code structure tree based on a mapping relation between software automation test elements and script code blocks, identifying a script range influenced by demand change, and outputting a change influence matrix; and based on the change influence matrix, generating an automatic script code snippet needing to be updated, and realizing preservation of the automatic script. The invention further provides a script preservation device and equipment, a storage medium and a program product.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Graphical user interface test case generation method based on multi-agent framework

The invention discloses a graphical user interface test case generation method based on a multi-agent framework, and relates to the field of graphical user interface tests.The graphical user interface test case generation method based on the multi-agent framework comprises the following steps that S1, a user demand document is obtained, and file text content and file format information are analyzed and extracted; s2, performing paragraph segmentation on the file text content to form logic semantic units; s3, carrying out correlation analysis on the segmented paragraphs, and obtaining a ternary file group; s4, performing paragraph annotation addition on the ternary file group, and constructing a semantic annotation result set; s5, extracting the agent framework to generate a graphical user interface test case, performing test recording, and storing a test result; s6, the result is verified, and a test case list and a software test report are generated. By conducting multi-dimensional analysis on the text content and the format structure in the document, the reduction efficiency from natural language description to structural semantic construction is improved.
Owner:BEIJING LANGUAGE AND CULTURE UNIVERSITY

Data storage method for artificial intelligence learning mode

The invention discloses a data storage method for an artificial intelligence learning mode, and relates to the technical field of computer data storage, and the method comprises the steps: 1, merging a multi-source perception stream into blocks in real time at the edge through monotone serial number writing, so as to provide a replayable time sequence; 2, asynchronous erasure coding is executed on the blocks, Merkel roots are calculated and written into a local cache, and dual guarantee of loss tolerance and integrity is achieved; 3, pushing slices and roots to object storage in sequence according to a network, and calling a time travel interface to solidify an incremental snapshot; 4, the cloud end monitors a snapshot hash event, a serial number chain is written through differential scanning, a gap is reconstructed through slices, an index is refreshed, and continuous consistency is kept; 5, the training process generates a Merkel proof online verification sample, and damaged data are immediately interpolated and repaired and an audit chain is recorded; and step 6, after training is finished, generating a leatherwise list and a frozen root, asynchronously cleaning redundant slices, updating a version table, and finally forming single-fingerprint traceable cost archiving.
Owner:北京爱宾果科技有限公司

Method and apparatus for clustering input data

A method comprising:obtaining a dataset of input samples, one of said input samples comprising a number of input data features,applying said input samples to a machine learning system comprising a first machine learning model, or encoder, configured to output encoded samples, one of said output encoded samples comprising fewer encoded features than the number of input data features,applying said output encoded samples to a second machine learning model, or decoder, of said machine learning system, configured to produce reconstructed input samples from said encoded samples,determining a reconstruction loss based on a difference between the input samples and the reconstructed samples,clustering said encoded samples into a plurality of clusters,determining a clustering error, said clustering error being defined as taking on a lower value the more homogeneous and separated the clusters are,obtaining a total loss based on the reconstruction loss and clustering error, andwhile a stopping condition comprising the total loss being less than a best total loss, is not reached, tuning internal weights of said encoder and said decoder based on said total loss, and reiterating the previous steps.
Owner:NOKIA SOLUTIONS & NETWORKS OY

Method and system for converting or encoding text

A publishing system with components, including:a system configured to receive at least one document including text that defines a base alphabet in one or more formats;a system configured to provide additional data for a reader to better understand the document which includes:a method of encoding or marking up non-phonetic words in the document to enable the reader to decode sounds of each non-phonetic word; anda system configured to output an encoded document with the text and the additional data in one or more formats,wherein the method of automatically encoding the non-phonetic words to make the encoded words phonetic:for at least one character (“spelling character”) in the non-phonetic word, using a compound character that includes the spelling character and a sound character, wherein the sound characters:are human-readable characters in the base alphabet and / or in one or more secondary alphabets,are added to the spelling characters to indicate that each spelling character makes the usual sound of the sound character,are added so that spelling characters can be visually discriminated from sound characters,are added such that a reader can recognize the non-phonetic word by sight because the spelling of the word is unchanged, andare added to the spelling characters such that the spelling characters and the sound characters remain human-readable such that the spelling character and the sound character of each compound character are within one visual field; andautomatically outputting the encoded words in a human-readable form / format such that the compound characters in the encoded word visually indicate which of the spelling characters have a sound other than their usual sound and what sound each character makes in the non-phonetic word when it does not make its usual sound.
Owner:STEPHEN CHRISTOPHER COLIN

Enterprise knowledge base retrieval and intelligent answering method and system based on large language model

The invention discloses an enterprise knowledge base retrieval and intelligent answering method and system based on a large language model. The method comprises the following steps: performing clause-level segmentation on an enterprise knowledge base document, associating document metadata to form structured knowledge entries, and establishing a keyword reverse index and a semantic vector index for the structured knowledge entries; analyzing the natural language query of the user, and performing multi-strategy expansion to generate an enhanced query expression and a query semantic vector; performing dual-channel mixed retrieval, performing duplicate removal, version filtering and weighted fusion sorting on a result, and generating a final candidate knowledge item list; and based on the candidate list and a predefined instruction, calling a large language model to generate a structured answer with complete traceability information. The method is compatible with an existing retrieval framework, precise understanding, knowledge point-level positioning, cross-document content integration and version consistency control of natural language problems are achieved, and the retrieval accuracy, answer availability and service intelligence level of an enterprise knowledge base are remarkably improved.
Owner:XIAMEN YUANTING INFORMATION TECH CO LTD

Multi-modal information analysis and scheme reminding method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a multi-modal information analysis and scheme reminding method, device, equipment and medium, which comprises the steps of receiving input data and converting the input data into multi-modal data, executing optical character recognition and image classification recognition to generate a recognition result, and sending the recognition result to a server; a natural language processing model is used for analyzing fuzzy description to generate an analysis result, a knowledge base is inquired, a knowledge graph is combined to generate an association result, the analysis result, the association result and user feature data are fused to generate an execution scheme, the execution scheme is compared with an abnormal list, supervision confirmation is triggered, and a compliance instruction is generated. And personalized reminding contents are generated. The information analysis integrity is improved through multi-modal recognition, natural language processing and the knowledge graph, supervision confirmation and personalized reminding are introduced, and intelligent, compliant and reliable reminding management is achieved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Automatic program repairing method based on data flow driving

PendingCN120872836AError detection/correctionValidation testData stream
The invention relates to an automatic program repairing method based on data flow driving, which takes an error file and a method signature as input, and reports error information in a defect positioning tool or an integrated development environment. Secondly, handing over error information and original defect codes to a large language model, extracting an element list related to defects from the large language model, then constructing defect contexts, using a search tool to retrieve key element definitions and dependency relationships, integrating control flow and error mode information, and finally obtaining defect information; then, the retrieved information, the original defect codes and the potential defect types are integrated into cue words, and the cue words are submitted to a large language model to generate candidate patches. And then the candidate patch is verified, and if the candidate patch does not pass the verification, error reporting information during the verification is combined with the candidate patch to carry out iterative optimization until the candidate patch successfully passes all test verification or the maximum number of iterations is reached. According to the method, context semantics are enhanced by means of data flow analysis, and the repairing accuracy and efficiency are improved.
Owner:CHONGQING UNIV

Intelligent composition quality evaluation method and system based on large language model

The invention relates to the technical field of artificial intelligence in the education industry, in particular to an intelligent composition quality evaluation method and system based on a large language model, and the method comprises the steps: carrying out the text normalization and semantic unit segmentation of a composition, extracting a semantic vector through a first large language model in combination with a context enhancement strategy, and positioning a semantic fracture risk position; recognizing composition core elements through a second large language model, and mapping the composition core elements back to the semantic unit sequence; constructing a demonstration logic diagram, extracting a core demonstration path and abstracting the core demonstration path into a logic role topological graph; in combination with a pre-constructed writing specification knowledge graph, comparing structural compliance, connection strength and an expected support relationship, identifying and demonstrating logic defects, and generating a global deduction item list; semantic clustering is carried out on illegal items to form an error label set, comprehensive weight is calculated in combination with historical data of students, and core weak items are positioned; according to the application, the logic analysis depth of intelligent evaluation of the argument is remarkably improved, and the pertinence and practicability of teaching feedback are improved.
Owner:DALIAN HOUREN EDUCATION TECH CO LTD

Business process standard dynamic optimization method and system based on large language model

The invention relates to a business process standard dynamic optimization method and system based on a large language model, and belongs to the technical field of computer application, and the method comprises the following steps: constructing a data collection strategy template; performing anomaly detection on the key field based on a preset three-level anomaly detection mechanism to obtain standard structured data; constructing a process execution accurate quantitative evaluation model, and constructing a process topological graph by using a process mining technology; calculating an abnormal probability of each node in the process topological graph by using a multi-task graph neural network, and obtaining a high-attention area in the process topological graph based on the abnormal probability; analyzing the multi-source heterogeneous data associated with the high-attention area through a large language model, positioning root causes, calculating root cause influence weights, and obtaining a root cause list; establishing an intelligent bidirectional feedback mechanism to generate optimization suggestions of the standard process; according to the method, the problems of serious data redundancy, insufficient standard dynamics and lack of feedback mechanisms in the traditional flow standard optimization technology in the power industry are effectively solved.
Owner:SICHUAN ZHONGDIAN AOSTAR INFORMATION TECHNOLOGIES CO LTD +3

Large language model (LLM)-based knowledge resource retriever and ranker

Disclosed herein are a system, method, and computer program product embodiments for retrieving and ranking knowledge resources relevant to a query from knowledge base(s). For example, a query for resources from knowledge base(s) may be received. Based on the query, a first set of candidate resources are obtained from the knowledge base(s) having a lexical similarity to the query search terms, and a second set of candidate resources are obtained from the knowledge base(s) having a semantical similarity to the search terms. For each of the first and second sets of candidate resources, a confidence level indicating the relevance of the candidate resource to the query is determined. The sets of candidate resources are ranked based on at least the confidence levels to generate a ranked list of candidate resources. A query response comprising at least a subset of the ranked list candidate resources is provided to a GUI.
Owner:SAP SE

Human resource intelligent management method and system based on man-post matching

The invention discloses a human resource intelligent management method and system based on man-post matching, and the method comprises the steps: receiving an unstructured post description text, extracting key information through a natural language processing technology, generating a structured multi-dimensional post portrait, and classifying the structured multi-dimensional post portrait; on the basis of historical recruitment data, predicting the number of future post demands by means of a time sequence analysis model; for the candidate resumes and the target post portraits, keyword correlation scores, depth semantic similarity scores and predictive stability scores are calculated in parallel; according to the post portrait classification application dynamic weight, performing weighted summation on the scores to generate a comprehensive matching score; and sorting the candidates according to the comprehensive matching scores, and outputting a sorted candidate list. According to the method and the system provided by the invention, the man-post matching accuracy and the recruitment efficiency are effectively improved, the core pain points of intelligent recruitment, man-post matching and demand prediction in human resource management are solved, and the method and the system are particularly suitable for the demands of human resource outsourcing and labor dispatching industries for stable service selection.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Translation ambiguity term accurate matching method based on fusion semantic vector space mapping

The invention discloses a fusion semantic vector space mapping-based translation ambiguity term accurate matching method, which comprises the following steps of: S1, obtaining source language ambiguity terms, context texts and a target language candidate translation list, and extracting domain tags and term matching features to form a multi-modal data set; s2, using improved XLM-R model coding to generate term-level, sentence-level and translation-level semantic vectors; s3, training a dynamic mapping matrix based on a bilingual parallel corpus, and aligning source side vectors to a shared semantic space; s4, fusing the source-side basic vector and the multi-dimensional features through a double-channel attention fusion network, and generating source-side and translation-side comprehensive semantic vectors; s5, introducing term-context attention weight to correct cosine similarity; and S6, outputting an optimal translation through normalized sorting and part-of-speech secondary judgment. According to the method, multi-field ambiguous term accurate matching is realized, the term translation precision and efficiency in professional fields are improved, and the requirements of high reliability of term translation in the fields of medicine, machinery, computers and the like are met.
Owner:XINJIANG DAWEIRAN BUILDING DECORATION GRP CO LTD

Large language model long text reasoning acceleration method and device based on speculative key value cache sparse technology, medium, terminal and program product

The invention provides a large language model long text reasoning acceleration method and device based on a speculative key value cache sparse technology, a medium, a terminal and a program product, the method is applied to electronic equipment comprising a GPU and a CPU, and the method comprises the steps that an index list of most important historical tokens is generated based on a distillation language model according to an obtained context sequence; comparing the index list generated at the current moment with the index list at the previous moment, and calculating to obtain a difference set part; asynchronously prefetching the key value cache of the difference set part from the CPU to the GPU; according to the asynchronously prefetched key value cache, performing parallel execution based on a large language model to generate a new token; obtaining a new sequence length according to the generated new token, and judging whether the new sequence length exceeds a preset threshold value or not; and if the threshold value is exceeded, executing unloading operation. According to the method, the performance and the stability of processing long text reasoning by the large language model can be improved, and key value cache optimization in the long context reasoning process is ensured to be always effective.
Owner:SHANGHAI JIAOTONG UNIV

Data acquisition method and equipment based on large language model, and medium

The invention relates to the technical field of electric digital data processing, in particular to a data acquisition method and device based on a large language model and a medium. The method comprises the steps of inputting a target statement into a large language model to obtain a target text output by the large language model, and obtaining a target keyword set of the target statement according to the target text; obtaining a target website list matched with the target keyword set from a preset website library; determining a target crawling time period of the website corresponding to each website and a crawling time step length corresponding to the target crawling time period according to a historical updating moment set of the website corresponding to each website in the target website list; and in the target crawling time period of the specified website, judging whether the data of the specified website is updated or not by taking the crawling time step length corresponding to the target crawling time period of the specified website as a judgment period, and if so, crawling the data from the specified website. According to the method, the data related to the statement input by the user can be comprehensively and effectively collected.
Owner:HANGZHOU YSCREDIT CO LTD

Whole-process application method of indoor decoration engineering digital model

The invention discloses a whole-process application method for an indoor decoration engineering digital model, and relates to the technical field of building information models.The method comprises the steps that a decoration semantic mapping engine is deployed, multi-source model attributes are standardized into a unified attribute dictionary, and a neutral semantic package is generated; splitting the neutral semantic package into single component nodes, and storing the single component nodes in a tense graph to realize object-level version management; generating an incremental packet by changing the monitoring stream and comparing the fingerprint; automatically combining decisions or generating a to-be-confirmed list by utilizing a conflict analysis pipeline; synchronizing the update fragment to a scheduling system and a logistics interface through a service bus; and finally, the field terminal dynamically superposes the scanning data and the sensing reading according to the component fingerprint, and writes back and feeds back to correct the mapping rule. The technical characteristics of the method cover semantic management, data synchronization and field feedback. The semantic consistency and the data accuracy are improved, the information islands are effectively inhibited, and the full-period management efficiency of the interior decoration engineering is improved.
Owner:GUANGZHOU QUANCHENG DUOWEI INFORMATION TECH CO LTD +1

Generating probabilistic data structures for lookup tables in computer memory for multi-token searching

Methods, systems, and non-transitory computer readable storage media are disclosed for optimizing computer memory usage for lookup lists in computer memory via probabilistic data structures. For example, the disclosed system generates a probabilistic data structure (e.g., a Bloom filter) to represent data in a lookup list including multi-token items by hashing items of the lookup list to sets of bit values in a bit vector. The disclosed system classifies text content in a digital document by utilizing a maximum number of tokens from multi-token items in the lookup list to select and compare sets of sequential tokens in the digital document to the probabilistic data structure. The disclosed system also iteratively reduces the number of tokens in sets of sequential tokens for subsequent comparisons. Furthermore, in some aspects, the disclosed system causes a computing device to modify a digital document and / or database operations based on the classifications.
Owner:ONETRUST LLC

Query method based on knowledge base and large model

The invention discloses a query method based on a knowledge base and a large model. The method comprises the following steps: constructing the knowledge base of a database; obtaining target language information which is input by a user and contains a user question; based on the knowledge base, generating a target pseudo mode corresponding to the target language information; based on the target pseudo mode, screening out a minimum table set associated with answering the user question from the knowledge base; packaging the multi-table connection logic in the minimum table set into a query view; based on the query view and the target language information, a reference example pair list which is most similar to the user problem and indicates the mapping relation between the problem and the query statement is retrieved from a knowledge base, and each example pair in the reference example pair list represents the mapping relation between the problem and the query statement; and generating a target query statement corresponding to the target language information based on the user problem, the target pseudo mode, the query view and the reference example pair list, and performing retrieval in a database based on the target query statement to obtain a query result.
Owner:JIUYOU TECH (SHENZHEN) CO LTD

Engineering knowledge base construction and deep retrieval method

The invention relates to the technical field of intelligent retrieval, in particular to an engineering knowledge base construction and deep retrieval method, which comprises the following steps of: in a knowledge base construction stage, performing knowledge level identification on an engineering knowledge document to form a multi-level tree structure taking chapter titles and corresponding contents as nodes; performing analysis processing on the heterogeneous content of each node to generate knowledge fragments; and generating embedded vectors of knowledge fragments by adopting a way of combining path embedding and content embedding, and storing association relationships among knowledge, vector data and original information to complete knowledge base construction. In the deep retrieval stage, after a user question is received, a knowledge blank question list is generated through query and rewriting, a knowledge base is retrieved through traversal of the list, the retrieval process is optimized through knowledge grouping, correlation sorting and loop termination judgment, and finally reply content subjected to traceable verification is generated and output based on existing knowledge accumulated in a circulation mode. According to the method, the reuse efficiency and the retrieval accuracy of the engineering knowledge can be improved.
Owner:CHINA TRANSPORT INFORMATION TECH GRP CO LTD