Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1645 results about "Sentence" patented technology

In non-functional linguistics, a sentence is a textual unit consisting of one or more words that are grammatically linked. In functional linguistics, a sentence is a unit of written texts delimited by graphological features such as upper case letters and markers such as periods, question marks, and exclamation marks. This notion contrasts with a curve, which is delimited by phonologic features such as pitch and loudness and markers such as pauses; and with a clause, which is a sequence of words that represents some process going on throughout time. This entry is mainly about sentence in its non-functional sense, though much work in functional linguistics is indirectly cited or considered such as the categories of Speech Act Theory.

Text prediction-based large-model real-time voice text intention recognition method and system

The invention discloses a large-model real-time voice text intention recognition method and system based on text prediction, and the method comprises the steps: obtaining the real-time voice data of a user, carrying out the real-time voice recognition processing through a streaming voice recognition interface, and obtaining a part of transcriptional text; inputting the partial transcription text into a mask language model for text prediction, and generating a plurality of high-credibility complete sentence candidates; based on the complete sentence candidates, the complete sentence candidates are input into a large language model in parallel for intention recognition, a corresponding intention result is obtained, and a mapping relation between the candidate sentences and the intention recognition result is established; and obtaining a sentence completely expressed by the user, calculating the similarity between the complete actual sentence and a plurality of high-credibility complete sentence candidates through a multi-level text similarity algorithm, selecting the candidate sentence with the highest similarity score, and directly obtaining a corresponding final intention recognition result based on the mapping relationship. The objective of the invention is to solve the technical problem of high response delay of an existing voice intention recognition system.
Owner:BEIJING YULORE INNOVATION TECH

Prompt generative model optimization system based on context

The invention relates to the technical field of natural language processing, in particular to a context-based Prompt generative model optimization system, which comprises a context analysis module, a cue word generation module, a context optimization module, a semantic check module and a structure reconstruction module. According to the method, the context path and the semantic hierarchy information of the semantic unit are introduced, the fine degree of semantic matching degree recognition is improved, semantic guide deviation caused by statement template solidification is avoided, the cue words are recombined in combination with the semantic coherence weight and the logic dependency relationship, and the recognition accuracy is improved. The consistency and expression accuracy of the prompt content in the context are enhanced, the prompt word insertion sequence and connection mode are dynamically adjusted through a semantic conflict detection and structure rechecking mechanism, coherence and stability of a semantic structure and controllable generation of the prompt content are kept, semantic conflicts and expression chaos caused by static matching are effectively avoided in the generation process, and the generation efficiency is improved. And dynamic adaptation of prompt configuration and smooth optimization of language output are integrally realized.
Owner:NALAI

Multimodal fusion entity retrieval enhancement generation method and device

The embodiment of the invention provides a multi-modal fusion entity retrieval enhancement generation method and device, and the method comprises the steps: carrying out the blocking and adaptive text extraction of multi-modal data in an offline stage, obtaining the text block data corresponding to each modal data, extracting the entity and relation of the text block data according to a language model, constructing an entity triple, and carrying out the segmentation and adaptive text extraction of the entity triple. Fusing the entity triad with the text block data to obtain an offline knowledge graph, and constructing a data index; in the present stage, a query statement of a user is received, text block data most similar to the query statement are retrieved in a knowledge graph through a data index, after entity aggregation is carried out on the retrieved text block data, the text block data are reordered according to retrieval scores, an entity aggregation result is obtained, entity ordering is carried out according to the entity aggregation result, and the entity aggregation result is obtained. By means of the multi-modal retrieval method and device, the efficiency and accuracy of multi-modal retrieval can be improved.
Owner:NO 15 INST OF CHINA ELECTRONICS TECH GRP

Translation ambiguity term accurate matching method based on fusion semantic vector space mapping

The invention discloses a fusion semantic vector space mapping-based translation ambiguity term accurate matching method, which comprises the following steps of: S1, obtaining source language ambiguity terms, context texts and a target language candidate translation list, and extracting domain tags and term matching features to form a multi-modal data set; s2, using improved XLM-R model coding to generate term-level, sentence-level and translation-level semantic vectors; s3, training a dynamic mapping matrix based on a bilingual parallel corpus, and aligning source side vectors to a shared semantic space; s4, fusing the source-side basic vector and the multi-dimensional features through a double-channel attention fusion network, and generating source-side and translation-side comprehensive semantic vectors; s5, introducing term-context attention weight to correct cosine similarity; and S6, outputting an optimal translation through normalized sorting and part-of-speech secondary judgment. According to the method, multi-field ambiguous term accurate matching is realized, the term translation precision and efficiency in professional fields are improved, and the requirements of high reliability of term translation in the fields of medicine, machinery, computers and the like are met.
Owner:XINJIANG DAWEIRAN BUILDING DECORATION GRP CO LTD

Enterprise knowledge graph automatic construction and intelligent retrieval method

The invention provides an enterprise knowledge graph automatic construction and intelligent retrieval method, which comprises the following steps: collecting multi-source data from a heterogeneous enterprise information system, and carrying out data cleaning and standardization processing; on the basis of a comprehensive scoring mechanism of field similarity and behavior semantic vectors, entities from different systems are merged, and a standard entity set with a unique identifier is generated; based on the standard entity set, in combination with a scoring mechanism of a task-type relationship and a collaborative relationship, extracting a semantic relationship from a behavior record, and constructing an enterprise knowledge graph structure; extracting representative semantic paths from the knowledge graph structure, screening high-quality paths through a path scoring model, and organizing the high-quality paths into a structured path index set; and receiving a natural language query statement, encoding the natural language query statement into a semantic vector, matching the semantic vector with the path index set, executing query in the atlas in combination with an authority control mechanism, and returning a result.
Owner:SHANGYANG TECH CO LTD

Natural statement decoding method and device based on high-density electrocorticogram

The invention discloses a natural statement decoding method and device based on high-density electrocorticogram. The method comprises the following steps: acquiring an electroencephalogram signal acquired based on the high-density electrocorticogram; different frequency band signals are extracted from the electroencephalogram signals; the signals of different frequency bands comprise high gamma frequency band signals and at least one frequency band signal with the frequency lower than that of the high gamma frequency band signals; acquiring a voice starting point corresponding to each frequency band signal, and determining a target voice starting point based on the voice starting point corresponding to each frequency band signal; after the target voice starting point, acquiring a syllable classification result and a tone decoding result corresponding to each frequency band signal, determining a target syllable classification result based on the syllable classification result corresponding to each frequency band signal, and determining a target tone decoding result based on the tone decoding result corresponding to each frequency band signal; and determining a target tone language corresponding to the electroencephalogram signal based on the target syllable classification result and the target tone decoding result. According to the scheme, the natural statement decoding accuracy can be improved.
Owner:SHANGHAI TECH UNIV +1

Knowledge fabric with mechanistic causal reasoning and deep language understanding

System and method for using knowledge fabric based on knowledge ontology, designed for deep language understanding and mechanistic causal reasoning, and meta-knowledge repository for auditable question answering. The method includes receiving an input text from a user, building a knowledge graph that represents real world facts and associations in the form of contextually tagged and weighted knowledge propositions, in multiple knowledge domains. The knowledge graph in combination with causal path knowledge and metadata describing digital sources containing answers constitutes the knowledge fabric. The method includes resolving ambiguity and determining actual intent of the user for the input text, from a plurality of interpretations of intent for sentences using the knowledge graph in conjunction with logical inference to achieve deep natural language understanding. The method includes finding / delivering response to the input request as to why / how unknown factors resulted in known outcome, or what outcomes are likely given known causal factors.
Owner:EMPATHI AI INC

Personalized Russian spoken language practice recommendation method and system based on artificial intelligence

The invention relates to the technical field of artificial intelligence education, in particular to a Russian spoken language practice personalized recommendation method and system based on artificial intelligence, and the method comprises the steps: 1, outputting a phoneme sequence with a timestamp through Russian automatic voice recognition; 2, collecting an exercise interruption position and repeated read-after behavior data; 3, generating a dynamic learner portrait; 4, mapping high-frequency errors in the learner portrait into abnormal path weights of map nodes; 5, a lattice tail error option and a non-matching body verb interference item are injected; 6, when the voice fluency attenuation of the learner exceeds a dynamic threshold value, the sentence complexity is reduced; and 7, calculating an error rate descent gradient based on the exercise completion data, and dynamically adjusting the abnormal path weight of the knowledge graph. Through audio stream analysis and syntax tree construction, the system can accurately identify errors of the learner in grammar, pronunciation and other aspects, and the learning efficiency is improved.
Owner:HARBIN UNIV

Text reinforcement learning method and device, electronic equipment and computer storage medium

The invention provides a text reinforcement learning method and device, electronic equipment and a computer storage medium, relates to the technical field of text reinforcement learning, and is applied to a text generation model. The method comprises the following steps: determining one or more target vocabularies in a statement to be analyzed, and generating a replacement word set according to the target vocabularies; replacing the target vocabulary in the to-be-analyzed statement with a current candidate replacement word to obtain a replacement statement; calculating one or more difference index values between the replacement statement and the to-be-analyzed statement; wherein the difference index value is used for quantifying the influence of the replacement of the target vocabulary on the original sentence from different dimensions; based on the difference index value, determining an importance weight of the target vocabulary in the to-be-analyzed statement; wherein the importance weight is used for redistributing the reward of the single character in the sentence. Vocabulary importance is quantified through disturbance analysis and multi-dimensional evaluation, and then fine-grained distribution of rewards is achieved.
Owner:CHENGDU HAPPY NOTE TECH CO LTD

A method, device, and medium for processing NOTAM text based on semantic enhancement

This invention relates to the field of text processing technology, and in particular to a method, device, and medium for processing navigational notice text based on semantic enhancement. The method includes: first, acquiring a content carrier to be processed; then, acquiring the semantic vector and glyph feature vector of the content carrier; concatenating the two types of vectors to form an enhanced text representation; extracting temporal features from the enhanced text representation to obtain temporal features containing forward and backward logical relationships within the text; acquiring the weights of words and sentences in the temporal features and performing weighting to obtain weighted word representations and weighted sentence representations; performing correction processing on the weighted representations to generate corrected text; and finally, validating the corrected text and outputting the target text. This invention can improve the accuracy and efficiency of content carrier processing.
Owner:CIVIL AVIATION UNIV OF CHINA

Intelligent error correction and style optimization system and method for business English writing

InactiveCN121145850ASemantic analysisPersonalizationGrammatical error
The invention discloses an intelligent error correction and style optimization system and method for business English writing, and relates to the technical field of natural language processing and artificial intelligence, and the system comprises the following components: a user interaction module, a data storage module, a language ability evaluation module and a self-adaptive error correction feedback and style optimization module. Through a multi-dimensional evaluation algorithm based on dynamic weight adjustment, historical writing data of a user is deeply analyzed, an evaluation system including four dimensions of grammar errors, vocabulary application, sentence structures and business profession is constructed, the system comprehensively evaluates the language ability of the user, and the user experience is improved. And the evaluation weight is dynamically adjusted according to the capability improvement rates of the users in different dimensions, so that the language capability levels of the users are more accurately divided, and the system can provide personalized error correction feedback and style optimization suggestions for the users with different language levels based on the evaluation result.
Owner:FUJIAN FORESTRY VOCATIONAL TECH COLLEGE

Text data processing method for English translation

The invention relates to the technical field of data processing, in particular to a text data processing method for English translation, which comprises the following steps of: preprocessing obtained English text data to be translated, and decomposing each sentence into a continuous vocabulary sequence; performing multi-level term recognition and extraction on the vocabulary sequence, and performing context semantic feature analysis on the English text data to be translated by using a deep language model; traversing each candidate term in the obtained initial candidate term set, performing term disambiguation in combination with the obtained context semantic feature set, and storing the term disambiguation in a constructed dynamic term library; and performing constraint translation on the sentence list obtained after preprocessing and the dynamic term library, and assembling the output translation sentences according to an original sequence to generate target language text data, thereby improving the accuracy of a translation result.
Owner:GUANGDONG OCEAN UNIVERSITY

Fine empirical tracing and declaration verification method based on heterogeneous target fusion

The invention relates to a natural language processing technology, in particular to a fine empirical tracing and declaration verification method based on heterogeneous target fusion, which comprises the following steps of: retrieving candidate documents related to a to-be-verified declaration, and extracting candidate sentences related to the to-be-verified declaration; constructing a cascaded statement-sentence association probability distribution model and a statement-sentence set association probability distribution model; learning prediction probability distribution between the declaration and a single candidate sentence based on a declaration-sentence association probability distribution model; screening a plurality of candidate sentences with relatively high empirical values to obtain a candidate empirical set; on the basis of a statement-sentence set association probability distribution model, capturing a semantic synergistic effect between candidate sentences in the candidate empirical set, and calculating global category probability distribution of the statement to be verified and the candidate empirical set; jointly optimizing parameters of the two models by adopting a heterogeneous target; and on the basis of the two trained models, fine empirical cause tracing and declaration verification are realized. According to the invention, a visual and accurate viewpoint verification result can be provided.
Owner:SUN YAT SEN UNIV

Anti-fact multi-mode dialogue emotion causal reasoning method based on double-branch hypergraph

The invention discloses an anti-fact multi-mode dialogue emotion causal reasoning method based on a double-branch hypergraph. The method comprises the following steps: respectively extracting sentence level feature vectors of three modes of text, voice and vision from input multi-mode dialogue data; carrying out modeling on a high-order relationship in the modals and between the modals by utilizing a hypergraph structure, and constructing a dialogue hypergraph containing multi-modal nodes and emotion nodes; introducing a hypergraph attention network on the hypergraph, learning contribution weight of each modal node to a target emotion node, and selecting a candidate reason node set; the candidate reason nodes are intervened, an anti-fact branch is constructed, a fact situation and final node feature representation under the anti-fact situation are calculated, and a causal effect vector is obtained; and designing a joint optimization objective function, and carrying out joint training on emotion recognition loss and causal consistency loss to realize synchronous prediction of emotion categories and emotion reasons. According to the method, a high-order semantic relationship can be effectively modeled in a multi-modal dialogue scene, and a key reason for emotion formation is reasoned.
Owner:JIANGSU UNIV

Knowledge graph construction method, device and equipment and readable storage medium

The invention discloses a knowledge graph construction method, device and equipment and a readable storage medium, and is applied to the technical field of natural language processing and knowledge engineering.The method comprises the steps that document content is divided to obtain initial document fragments, and all the initial document fragments are merged and divided based on semantic similarity to obtain target division blocks; based on the target division block, subject-predicate-object formatting processing is carried out to obtain a subject-predicate-object formatting result; entity and relation extraction is carried out according to the subject-predicate-object formatting result to obtain a display triple, implicit relation reasoning is carried out to obtain an implicit relation triple, entity and relation type normalization is carried out to obtain a normalized triple, and the knowledge graph is constructed based on the normalized triple. The block segmentation driven by semantic similarity is adopted to avoid sentence breakage, cascade errors are reduced based on subject-object formatting processing, and the problems of fragmentation, link missing and the like are solved based on implicit relations, so that the integrity and accuracy of knowledge graph construction are improved.
Owner:SICHUAN SHUTIANMENGTU DATA TECH CO LTD

SEO content detection and release gating method and system based on information gain

The invention discloses an SEO content detection and release gating method and system based on information gain, and the method comprises the steps: carrying out the paragraph-level viewpoint sentence decomposition of a collected search engine result page document and a candidate document for a target query, and constructing a consensus semantic field; performing semantic alignment and screening on the viewpoint sentences of the candidate documents and the consensus semantic field, and performing multi-source evidence verification and reliability aggregation on the screened effective newly-added viewpoint sentences to obtain macroscopic semantic novelty measurement of the candidate documents; constructing a consensus knowledge graph to identify the structural hole and calculating the gain of the structural hole; and carrying out joint scoring and issuing decisions, and outputting a structured evidence packet. According to the method, evaluation of novelty and information gain is quantified by constructing a ranking-weighted consensus semantic field; the reliability of the measurement gain is ensured through evidence conditional macroscopic divergence calculation and viewpoint sentence-level multi-source verification; and a structured evidence packet organized according to the viewpoint sentences is generated, so that the decision is well documented and auditable.
Owner:TOUCHDATA

Large model effect evaluation method and device, storage medium and computer equipment

The invention provides a large model effect evaluation method and device, a storage medium and computer equipment. Specifically, the to-be-evaluated large model is a large language model constructed according to the target prompt statement. According to the method, the to-be-evaluated text can be generated according to the session reply text of the to-be-evaluated large model for the at least one to-be-evaluated data set, and the to-be-evaluated text is evaluated by using the target judgment model corresponding to the target evaluation index, so that the effect evaluation result of the to-be-evaluated large model is determined. The to-be-evaluated text is a comprehensive result of combined action of the large language model, the target prompt statement and the to-be-evaluated data set, so that the method can cover three perspectives of the large model, the prompt statement and the data set at the same time, multi-angle and comprehensive evaluation is realized, and the effect and the service capability of the large model can be accurately obtained. Moreover, automatic scoring is carried out by adopting the target judgment model, so that the evaluation subjectivity and the labor cost can be greatly reduced, and large-model evaluation can be efficiently realized.
Owner:GUANGZHOU QUYAN NETWORK TECH CO LTD

Anaphora disambiguation method and system based on big language model enhanced text and structured query language generation

The invention discloses an anaphora disambiguation method and system based on big language model enhanced text and structured query language generation. The method comprises the following steps: receiving a natural language question of a user, and analyzing and generating structured table field information through mode information; identifying and rewriting fuzzy time expression in the problem; extracting ambiguous entities, and replacing the ambiguous entities with database standard values through semantic matching and fuzzy matching; synthesizing the information to generate a structured query language statement and executing the structured query language statement; and if the execution result is null, automatically triggering an anaphora disambiguation process, performing error correction and re-matching on field values in the statement through a dual matching mechanism, and generating and executing a corrected query statement. According to the method, the problems of fuzzy anaphora, indefinite time expression, low matching accuracy, resource waste and the like in a traditional SQL system are effectively solved through a strategy of combining pre-processing and post-processing, and the accuracy, robustness and execution efficiency of complex query are remarkably improved.
Owner:CHINA TELECOM DIGITAL INTELLIGENCE TECH CO LTD

Public opinion information detection method, device and equipment based on heterogeneous large model

The invention relates to the field of network public opinions, and discloses a public opinion information detection method based on a heterogeneous large model, which comprises the following steps: acquiring a public opinion text, performing sentence segmentation, denoising and word segmentation preprocessing, and inputting the processed text into a detection model to output a harmful information category and an early warning level. The detection model is composed of a first large language model and a second large language model, and has the structural characteristics of cross-architecture semantic alignment, hierarchical knowledge distillation, field attention enhancement, multi-channel decision fusion and the like. Wherein the cross-architecture semantic alignment realizes hidden space sharing through bidirectional projection; the hierarchical knowledge distillation dynamically distributes weights according to task contribution of each layer; introducing domain bias to enhance semantic focusing by domain-enhanced attention; and the output of the two models is adaptively integrated and classified through multi-channel decision fusion. According to the method, the accuracy, robustness and reasoning efficiency of public opinion harmful information detection are effectively improved.
Owner:BEIJING ZHIHUI XINGGUANG INFORMATION TECH CO LTD

Domain-aware autocomplete

Various embodiments of the present disclosure provide model-based domain-aware autocomplete techniques for generating autocomplete suggestions in a complex search domain. Example embodiments are configured to generate, using a domain-aware autocomplete model, a label for an autocomplete suggestion based on a set of keywords within an autocomplete suggestion training dataset associated with a target domain source. Example embodiments are also configured to generate, using a weak-labeling model, an updated label for the autocomplete suggestion by decorrelating the set of keywords from the label. Example embodiments are also configured to generate, using a sentence classification model, a category for the autocomplete suggestion based on the updated label. Example embodiments are also configured to, using the domain-aware autocomplete model, generate a suggestion-category pair (SCP) based on the autocomplete suggestion and the category for the autocomplete suggestion. Example embodiments are also configured for initiating performance of a search query resolution based on the SCP.
Owner:OPTUM INC

RAG application-oriented context poisoning attack defense method

The invention discloses a context poisoning attack defense method oriented to an RAG application, and relates to the technical field of RAG. the method comprises the following steps: inputting a target query statement, and retrieving the target query statement to obtain multiple pieces of context information; taking representative sentences in the retrieved context information, and identifying and filtering potential malicious template clusters; the big language model gives all candidate answers according to existing context information, the logarithmic probability of all contexts to different candidate answers is calculated, and after the influence of parameter knowledge of the big language model is removed from the logarithmic probability, the support degree of all contexts to different candidate answers is obtained; the whole logarithmic probability vector is used as a support degree distribution condition of the context to the candidate answers; identifying a single piece of harmful information from the support degree distribution condition of the context to the candidate answers through a logistic regression model so as to filter wrong answers; according to the attack defense method provided by the invention, centralized injection of multiple malicious texts and sparse injection of a small number of malicious texts can be defended.
Owner:SOUTHWEST PETROLEUM UNIV

Unstructured long text question and answer method and system based on logic map enhancement

The invention relates to an unstructured long text question and answer method and system based on logic map enhancement, and a semantic map enhanced question and answer mechanism is formed by combining semantic block clustering and logic relation modeling. The system encodes paragraphs through Sension-BERT, identifies multi-level semantic blocks based on a Louvain algorithm, performs multi-type logical relationship identification on paired semantic blocks through a trained logical relationship classifier (such as model fine adjustment based on RoBERTa-large and the like), constructs a structured logical map including causal, support, comparison and other relationships, and performs multi-type logical relationship identification on the semantic blocks. Finally, a structured and high-interpretation answer is generated by combining graph representation and context input large language model, the long-distance dependency understanding and causal chain generation capability in a long text is remarkably improved, and the method is suitable for complex scenes such as annual report analysis, policy interpretation and scientific research literature.
Owner:SHANGHAI ACADEMY OF SOCIAL SCIENCES +1

Standard document authoring method and system enhanced based on knowledge retrieval

The application discloses a standard document writing method and system based on knowledge retrieval enhancement, comprising the following steps: acquiring field multi-source heterogeneous knowledge by knowledge retrieval and constructing a field knowledge graph and a content increment model; performing semantic extraction, structure division, content segmentation and text granularity perception; determining first matching information and second matching information; matching standard paragraph structures and standard sentence structures to generate a standard template; performing compliance review to obtain a compliant text and recommended text; performing abnormality screening and abnormality correction to obtain a compliance value; inputting the compliant text, the recommended text and the compliance value into the content increment model to perform normalized conversion to obtain standard materials; and inputting the standard materials of each sentence into the standard template to obtain a standard document. The method can improve the efficiency and accuracy of standard document writing, has good interpretability, and can be directly applied to a standard document writing system based on knowledge retrieval enhancement.
Owner:CHINA STANDARD TECH DEV CORP

Method for bidirectional translation between sign language and text using ai, deep learning, and dictionary search techniques

The present invention facilitates communication between sign language users and machines by translating sign language and text using AI models, deep learning computer vision, and word embeddings. Users interact via sign language, captured and processed through deep learning and NLP modules. The system converts sign language videos into text, constructs coherent sentences, and generates contextually appropriate responses using a Retrieve and Generate (RAG) model. Responses are translated back into sign language videos, spelling out words not found in the dictionary. If requested, a human agent can respond. Key features include high-accuracy recognition, context-aware response generation, dynamic vocabulary updates, and optional human interaction. The method ensures efficient processing with LLM, embedding techniques, and deep learning, optimizing translation accuracy and user experience. The system adapts to multiple languages and dialects by training on specific sign languages, making it applicable globally.
Owner:MAHGOUB AHMED

Code similarity detection method and system based on large language model

The invention discloses a code similarity detection method and system based on a large language model, and the method comprises the steps: obtaining a to-be-detected source code pair, and marking the to-be-detected source code pair as a source code A and a source code B; performing code cleaning and format standardization on the source code A and the source code B, and mapping a variable name and a function name which are customized by a user into a uniform placeholder; analyzing the source codes A and B based on the abstract syntax tree, respectively replacing variable names and function names in the source codes A and B with unified serialized placeholders, and maintaining a mapping table; meanwhile, expanding a lexical dictionary of the pre-training large language model, and inserting a special identifier; splicing the replaced code snippets with special identifiers, and constructing a structure sensing input sequence; inputting the structure perception input sequence into a pre-trained large language model backbone network for feature coding to obtain a high-dimensional semantic feature vector containing global context information; connecting a multi-task prediction head behind the large language model backbone network, inputting the feature vector into the multi-task prediction head, outputting probability distribution of code clone types through a classification task head, and respectively outputting a row level similarity score and a lexical element level similarity score through a regression task head; and according to the classification probability and the regression score, performing comprehensive judgment by combining a preset threshold, and generating a detection report. According to the method, similar codes after variable renaming, statement rearrangement or control flow transformation can be accurately recognized, and the accuracy and robustness of code similarity detection are improved.
Owner:NANJING UNIV OF SCI & TECH

Searching method and device, storage medium and program product

The invention relates to the technical field of artificial intelligence, and discloses a search method and device, a storage medium and a program product. The method comprises the steps of obtaining a query statement input by a user, and performing semantic analysis on the query statement to obtain an intention keyword; according to the intention keyword, searching in a knowledge graph to obtain an initial sub-graph, and according to the initial sub-graph and a preset statement template, obtaining a standard query statement; and through a pre-trained large language model, according to the initial sub-graph and the standard query statement, obtaining a multi-modal search result. According to the scheme of the embodiment, the problem search is performed through the preset standard statement template in combination with the knowledge graph and the large language model, so that the defects of the large model in logical reasoning and knowledge coverage can be overcome, and the accuracy of the search result can be improved.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Content abstract generation method based on chapter structure analysis

The invention discloses a content abstract generation method based on chapter structure analysis, and belongs to the technical field of natural language processing. The method comprises the steps that firstly, an original text is preprocessed, then text structure deep analysis is carried out, the text type of the text is recognized, an explicit / implicit text relation is extracted, and special symbols are introduced through a Prompt normal form for implicit text relation extraction to strengthen logic semantics; secondly, scoring sentences by adopting a double-path scoring mechanism in combination with a deep neural network model of chapter structure features and an optimized text sorting algorithm, and fusing scores through a logistic regression model; then, screening target sentences based on a chapter relation weighted secondary modulus function and a greedy algorithm, and finally, carrying out post-processing to generate an abstract. According to the abstract generation method, the chapter structure logic is deeply utilized, so that the problems of logic unsmoothness, information redundancy or key relation missing in the existing abstract generation are solved, and the semantic coherence and information integrity of the abstract are improved.
Owner:MAIGET INFORMATION TECH (BEIJING) CO LTD

Entity recognition-based collaborative problem traceability and strategy matching method

The invention discloses a collaborative problem traceability and strategy matching method based on entity recognition, and the method comprises the steps: carrying out the semantic expansion of a negative problem text through a T5 generation type pre-training model, and extracting time, a subject, a problem, a strategy and other entities and position information through a BERT-BiLSTM-CRF model; splicing the subject entity, the question and the strategy entity to generate a text with a prefix, correcting by using a BART generation type pre-training model, and classifying and mapping the question to a subsystem for storage through a BERT-DPCNN model; coarse screening is conducted on texts in the subsystem through a regular rule base, vector representation is generated in combination with BERT and RoBERTa models, cosine similarity is calculated, and matching results are determined by synthesizing sentence positions and keyword overlapping degrees; constructing a question-strategy keyword knowledge graph based on a matching result, and inputting the keywords into a DeepSeek-V3 big language model to generate an optimization strategy scheme; according to the method, multiple deep learning models and large language models are fused, so that semantic understanding, entity extraction and accurate matching of negative problem texts are realized, and a final optimization strategy scheme is generated.
Owner:TIANJIN UNIV

Simultaneous interpretation method and system based on large model and electronic equipment

The invention discloses a simultaneous interpretation method and system based on a large model and electronic equipment, and the method comprises the steps: extracting bilingual parallel corpora related to terms from professional resources based on a standardized professional dictionary, obtaining qualified corpora through data enhancement processing and manual screening, and constructing a multi-level corpus according to the levels of words, sentences and paragraphs; the method comprises the following steps: receiving an input audio stream in real time, extracting acoustic features through preprocessing, inputting a pre-established large-scale speech recognition model, and carrying out incremental decoding on the acoustic features in a sliding window mode; and calling a sentence boundary prediction network to judge a pause point, and outputting a text stream with a timestamp. Performing fine tuning on the large-scale speech recognition model by using a multi-level corpus, translating a text stream based on the fine-tuned large-scale speech recognition model, and constraining term translation according to a standardized professional dictionary; and synchronously displaying the audio output in the translation result and the subtitles. According to the scheme, the terminology recognition and translation accuracy is improved, and simultaneous interpretation delay is reduced.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Machine learning-based literature search and retrieval and related machine learning model training methods

Described herein are systems and methods for performing literature retrieval and related machine learning model training methods. An example computer-implemented method of training a machine learning model configured for literature retrieval I includes receiving a plurality of full-text articles; extracting, from the plurality of full-text articles, a plurality of positive sentence-citation pairs, each positive sentence-citation pair comprising a respective citing sentence and at least one cited article that is associated with the respective citing sentence; creating a labeled dataset comprising the plurality of positive sentence-citation pairs; and training a machine learning model using the labeled dataset.
Owner:FLORIDA STATE UNIV RES FOUND INC