Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

215 results about "Headword" patented technology

A headword, head word, lemma, or sometimes catchword, is the word under which a set of related dictionary or encyclopaedia entries appears. The headword is used to locate the entry, and dictates its alphabetical position. Depending on the size and nature of the dictionary or encyclopedia, the entry may include alternative meanings of the word, its etymology, pronunciation and inflections, compound words or phrases that contain the headword, and encyclopedic information about the concepts represented by the word.

Semantic enhancement adaptive partitioning method and system for natural resource large model questions and answers

The invention provides a semantic enhancement adaptive partitioning method and system for natural resource large model questions and answers, and aims to solve the problems of difficulty in term boundary recognition, damage to semantic integrity and the like. According to the method, three core technologies including theme perception coarse-grained paragraph division, self-adaptive sliding window theme hierarchy division and embedded perception context self-adaptive text segmentation are fused. The system firstly analyzes a natural resource long text structure, identifies titles and theme levels and aligns associated contents; paragraphs are extracted according to a theme perception strategy and are subdivided into sentence sets according to grammar rules; an improved sliding window mechanism is adopted to divide sentences into window sentence block groups. The method is characterized in that a dynamic aggregation threshold mechanism is introduced, the semantic association degree between adjacent sentence blocks is calculated through an embedded perception context semantic segmentation technology, whether the sentence blocks are combined or not is judged by combining a similarity distribution change trend and a dynamic adjustment threshold, self-adaptive delimitation of semantic boundaries is achieved, and text blocks which are clear in structure and coherent in semantics are generated.
Owner:HUBEI PROVINCIAL DEPT OF NATURAL RESOURCES INFORMATION CENT +1

Document content extraction method and system based on multimodal model collaboration, terminal and medium

The invention belongs to the technical field of document content extraction, and particularly discloses a document content extraction method and system based on multimodal model collaboration, a terminal and a medium. Comprising the following steps: identifying the type of an input to-be-processed document, and judging the document type; on the basis of the type identification result, calling a multi-modal model to analyze the document content, and outputting space coordinates, visual features and semantic features of document elements; generating a content sequence according with a reading habit through a semantic sequence reconstruction algorithm; paragraph boundary detection, paragraph recombination and semantic association modeling of charts and texts are completed based on the multilayer attention network and the graph neural network; grammar error correction, format optimization and title hierarchy generation are carried out by using a large language model and a hierarchical classification network; and converting the identification result into a structured output file. According to the method, the processing requirements of different types of documents can be considered, and high-precision analysis and efficient output are realized under the scenes of complex layouts, multiple languages and formula tables.
Owner:TUOSI (SHANDONG) INFORMATION TECHNOLOGY CO LTD

Highway intelligent operation and maintenance question-answering system based on large language model

The invention provides a highway intelligent operation and maintenance question-answering system based on a large language model, and belongs to the technical field of natural language processing. The system takes a large language model as a core reasoning engine and combines a domain knowledge base and an RAG technology to realize accurate question and answer of highway operation and maintenance; the method comprises the following steps: based on original knowledge data cutting, generating a title through a large language model, and customizing a knowledge base; receiving query, analyzing an intention by using a large language model, and matching to generate a function; the query is rewritten by using a large language model, dense and sparse vector query is generated, and a double-layer retrieval mechanism is formed; a two-step recall mode is utilized, coarse-grained recall is firstly carried out, then a recall result is subjected to fine-grained optimization through a screening mechanism, and a reasoning text is generated; and finally, inputting the query and reasoning text into the large language model, and generating an optimal answer through single-round and multi-round questions and answers. According to the method, the professionality and reliability of answers are enhanced, and the technical problem that answers are incomplete and inaccurate in an existing question and answer system is solved.
Owner:KUNMING UNIV OF SCI & TECH

Method and system for retrieving DOCX document content based on keywords

The invention belongs to the technical field of text processing, and particularly relates to a method and system for retrieving DOCX document content based on keywords, which comprises the following steps: analyzing an Office Open XML structure of a DOCX document, combining with multi-dimensional features such as style names, and utilizing a title classification score model to accurately distinguish a title and a text, so that a semantic hierarchical structure of the document is effectively reserved; and secondly, a multi-level semantic extension mechanism is introduced, and a Sension-BERT, a HowNet knowledge base and a Word2Vec model are fused, so that intelligent extension of synonyms and synonyms of keywords is realized, and the recall rate and semantic understanding ability of retrieval are remarkably improved. And in addition, a BM25 model is combined with paragraph length normalization and structure position weight to calculate a correlation score, so that retrieval results are sorted more accurately and reasonably. The construction of the reverse index is combined with the position coding and compression optimization strategy, and the retrieval efficiency and the storage performance are both considered.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Vehicle knowledge question-answering method based on GraphRAG method and computer equipment

The invention discloses a vehicle knowledge question-answering method based on a GraphRAG method and computer equipment, and the method comprises the steps: building a document tree structure based on the title hierarchy of a vehicle knowledge document, distributing a text identifier, and adding metadata information; converting the stored text units into text unit vectors by adopting a vector embedding model, and establishing a text vector index; calculating a similarity score between the query vector and the text unit vector, and determining a traditional RAG path text unit based on the similarity score; identifying entities and entity relationships of the text units according to the entity types and the relationship types by adopting a large language model so as to generate a vehicle knowledge graph; performing multi-level community division on the vehicle knowledge graph by adopting a Leiton community detection algorithm, and generating a structured community abstract according to each level of community; determining a GraphRAG path community abstract by adopting a large language model; and fusing the community abstract of the GraphRAG path with the text unit of the traditional RAG path and generating an answer.
Owner:FAW VOLKSWAGEN AUTOMOTIVE CO LTD

Structured text analysis method and system based on title recognition and hierarchical abstract

The invention relates to a structured text analysis method and system based on title recognition and hierarchical abstracts, and the method comprises the steps: receiving a text, and obtaining sample data of a training language model; the vector database can perform vector dimension adjustment on the data of the stored text. The vector database takes vectors as basic storage units, and converts unstructured data into high-dimensional vectors through an embedding technology. And calling the data storage type information of the title text data set to determine the data storage type of the title text data set matched with the text storage mode. And obtaining a structure type of sample data of each text by utilizing a semantic segmentation script. And segmenting the document according to the identified chapter titles, and subdividing each chapter into a plurality of logic paragraphs to form structured text block units. The multi-level abstract generation model matched with the target model can fully mine level information and semantic association in the document. Therefore, the effect of analyzing the structured text by the intelligent question-answering system is improved.
Owner:HANGZHOU MEITENG TECH CO LTD

Water conservancy knowledge structured extraction and verification method and device

The invention provides a water conservancy knowledge structured extraction and verification method and device, and belongs to the technical field of artificial intelligence, and the method comprises the steps: carrying out the differential text processing of different formats of files, and generating an intermediate file; classifying the intermediate file into a regulation class or a non-regulation class based on a preset rule base; performing hierarchical title identification on the regulatory files to form entry knowledge blocks, and converting table contents into HTML (Hypertext Markup Language) knowledge blocks; performing semantic segmentation on the non-regulation file to generate knowledge blocks; performing knowledge block checking and filing, and marking an abnormal alarm block; converting the table knowledge blocks into natural language description by utilizing a large model; and positioning the context of the original text of the alarm knowledge block, and performing intelligent correction through a large model. According to the method, a traditional semantic analysis model and a large language model are creatively fused, a closed-loop process of preprocessing, extraction, verification and correction is formed, the problems of structured analysis and error correction of complex texts in the water conservancy field are solved, and the knowledge processing efficiency and accuracy are remarkably improved.
Owner:长江水利委员会网络与信息中心

Semantic understanding-based environmental impact report auxiliary auditing system

The invention relates to the technical field of semantic understanding, in particular to an environmental impact report auxiliary auditing system based on semantic understanding, which comprises a text cleaning module, a chapter segmentation module, a rule generation module, a content auditing module and a result output module. According to the method, the paragraph structure and the layout identification are analyzed in a unified mode, key paragraph types in environmental influence chapters are efficiently recognized, the text integration accuracy is improved by combining inter-paragraph similarity calculation and repetition rate screening, chapter boundary accurate positioning and affiliation adjustment are achieved on the basis of format feature comparison of serial numbers and title styles, and the text integration efficiency is improved. A multi-layer parameter alignment rule set is constructed to support comprehensive verification of standard numbers, monitoring frequencies and periodic elements, an exception labeling mechanism is combined to complete difference item logic judgment and consistency evaluation, an auditing basis with rule driving and data comparison capabilities is formed, chapter positioning precision, element extraction comprehensiveness and logic exception recognition efficiency are improved, and the method is suitable for popularization and application. And the pertinence and the automatic processing depth in the auditing process are enhanced.
Owner:SHANGHAI RUIDUN INFORMATION TECHNOLOGY CO LTD +1

Knowledge base construction and retrieval method and system based on multi-source text in building field

The invention relates to the technical field of building information, and provides a knowledge base construction and retrieval method and system based on a multi-source text in the building field, and the method comprises the following steps: a knowledge base construction stage: constructing a multi-dimensional metadata feature vector for a multivariate text based on a standard classification index table; the method comprises the following steps of: converting and segmenting a document, splicing an end clause with all superior title texts by utilizing a context inheritance algorithm to form a text unit with complete semantics, and performing dynamic filtering based on an analyzed query intention and a metadata vector at a user retrieval stage; then, in the screening set, performing fusion calculation on semantic vector similarity, keyword matching degree and authority offset weight based on effectiveness attribute and implementation time, and performing mixed retrieval and reordering on the text units; and finally selecting a text unit according to a sorting result and inputting the text unit into the large language model to generate answers. According to the method, high-precision and high-compliance intelligent retrieval and question answering of building domain knowledge are realized.
Owner:SHANGHAI RESEARCH INSTITUTE OF BUILDING SCIENCES CO LTD

Table text boundary adaptive information extraction method and system

The invention provides a form text boundary adaptive information extraction method and system in the technical field of computer information processing. The method comprises the following steps: S1, identifying a title frame from a form image; s2, calculating the aspect ratio of the title frame to determine the title arrangement direction of the table titles; s3, detecting a text detection box of the table image, and performing classification prediction of a text arrangement direction on the text detection box by combining a text direction model with a title arrangement direction; s4, controlling affine transformation of the table image through the text arrangement direction and the text detection box to obtain a corrected image; s5, identifying a table title in the corrected image, and matching a table type from the table type configuration file based on the table title; s6, extracting keywords from the corrected image based on the table type; and S7, processing the keyword to extract table information. The method has the advantages that the adaptability, accuracy and efficiency of table information extraction are greatly improved.
Owner:FUJIAN NEWLAND SOFTWARE ENGINEERING CO LTD

System and method for creating a controllable output summary from text

A system for creating a controllable output summary of text is disclosed. The system generates a set of summaries of text and for a first summary from among the set of summaries, executes a script that is configured to append the first summary to a set of summary-title pairs. The system generates a first title associated with the first summary in response to executing the script. The system compares the first title with the text. Based at least on the comparison, the system determines if the title indicates the context of the text. If it is determined that the title indicates the context of the text, the system generates a dataset of title-text pairs including the first title paired with the text. The system trains a target summarization algorithm with the generated dataset.
Owner:BANK OF AMERICA CORP

Knowledge base context awareness and traceability enhanced intelligent retrieval and question-answering system

The invention provides an intelligent retrieval and question-answering system for knowledge base context awareness and traceability enhancement, and the system obtains an analysis result, an optical character recognition result and a semantic analysis result through the analysis of PDF physical layout, Word / Excel paragraphs, titles, tables and style attributes thereof, and the optical character recognition. Structured knowledge blocks, positions and levels of path images and vector representation are generated and stored in a vector database, and related results of user query requests are extracted by executing a mixed retrieval algorithm with keyword retrieval and vector semantic retrieval. Performing intelligent reordering by considering semantic similarity, keyword matching, knowledge block type weight, source knowledge base weight and page position weight to obtain a candidate knowledge block list, constructing cue words according to an intelligent reordering result, and calling an external large language model to generate answers; source labels in answers are managed to be associated with metadata of corresponding numbers in a candidate knowledge block list, and the problem that text blocks and traceability are not accurate is solved.
Owner:CHINA HAISUM ENG

Duplication check system and method for paper generated by artificial intelligence

A duplication check system and method for paper generated by artificial intelligence includes steps: S1: the user uploading the academic paper to be detected to a system, and the system automatically extracting the title, the abstract, and the headline of each paragraph of the paper; S2: fusing the title, the abstract, and the headline of each paragraph of the paper with the contextual information of the paper and extracting theme features; S3: after the different themes of the paper are extracted, repeatedly using similar tones for each theme in all different AI tools to propose integrate text requirements, searching each theme for times of the number of repetitions of integration in each AI tool until no new content is obtained, matching all the obtained texts with the paper to be duplication checked, based on natural language understanding, for duplication check, and marking the matching repeated parts and indicating the sources.
Owner:WU JIANG

Cross-modal fusion lightweight defect detection method based on knowledge distillation

The invention belongs to the technical field of digital image processing, and particularly relates to a knowledge distillation-based cross-modal fusion lightweight defect detection method, which comprises the following steps of S10, cross-modal fusion distillation; through a bidirectional vision-language alignment mechanism, the frozen multi-modal knowledge of a teacher model vision-language basic model VLM is migrated to a lightweight student model, and the dual-path fusion module comprises text condition region representation injected with semantic context and region anchoring semantic embedding fused with spatial vision clues; step S20, cross-header word-region alignment is carried out; embedding the fusion visual features generated by the two-way fusion module and the enhanced text to generate cross-head prediction so as to simulate the semantic-space association capability of a teacher model; s30, knowledge distillation loss is fused; according to the method, the multi-modal basic model is fused and distilled into the lightweight single-modal detection model, the detection performance in a defect detection scene can be improved, and compared with the basic model, the reasoning speed is greatly improved, and the parameter quantity is reduced.
Owner:CENT SOUTH UNIV

Sparse N-gram modeling for patient-entity relation extraction

Methods, systems, and software are provided for determining a relationship between a subject and a health entity. An electronic health record (EHR) for the subject is split into sections by detecting delineating section headers, and sections are subdivided into text spans. Text spans are filtered by language pattern recognition into a set of text spans having an expression related to the health entity. The natural language context of the expression in each text span in the set is evaluated to obtain a corresponding scoring representation. Scoring representations are inputted into a model comprising a plurality of parameters. The model outputs, for each text span in the set, at least a prediction that the text span is associated with the health entity. Models for determining relationships between subjects and health entities and methods for training models to determine relationships between subjects and health entities are also provided.
Owner:TEMPUS AI INC

Book-based subject knowledge system construction method and device

The invention discloses a book-based subject knowledge system construction method and device, and the method comprises the steps: representing a directory as a directed tree, and dividing the directed tree into a skeleton module and a plurality of chapter modules; carrying out the association and expansion of a title through a large language model, and generating an abstract; knowledge concepts, subject terms and structural relations in titles in all the modules are extracted and subjected to double correction to obtain a plurality of knowledge system trees, and then a mother tree obtained after the skeleton module is processed and a sub-tree obtained by the chapter module are spliced; and finally, giving a plurality of knowledge trees obtained according to a plurality of groups of book directories in the same subject field, and combining the knowledge trees into a unified knowledge system tree based on a finite-state machine principle. According to the method, the knowledge system is constructed and decoupled into an extraction stage and a fusion stage based on LLM, the problem of a pain point of drifting between knowledge concepts and attributes when a simple chapter title is faced is solved, and strict control on semantic consistency and hierarchical rationality in a path merging process is ensured based on a state machine.
Owner:ZHEJIANG UNIV

Method and device for intelligent semantic error correction and business term optimization of foreign trade letter electricity

The invention relates to the technical field of natural language processing, in particular to a foreign trade letter intelligent semantic error correction and business term optimization method and device, and the method comprises the steps: obtaining a target foreign trade letter, and constructing a target corpus; performing Chinese word segmentation and part-of-speech tagging on the target foreign trade letter; performing term optimization based on the word segmentation result and the knowledge graph; identifying the letter title by using a conditional random field model, and converting the letter title into structured data; a Bi-LSTM-CRF model is adopted to carry out risk point detection, including Bi-LSTM coding, feature engineering and Max-pooling technologies, a part-of-speech sequence is obtained through a softmax function and a Viterbi path, and sequence labeling is carried out to obtain a risk point detection result; and finally, performing Chinese error correction based on the word segmentation result after part-of-speech tagging and the knowledge graph. The recognition and correction accuracy of foreign trade terminologies is improved, and communication obstacles caused by nonstandard use of the terminologies are effectively reduced.
Owner:GUANGDONG VOCATIONAL COLLEGE OF SCI & TRADE

Document segmentation method and device based on large language model, equipment and storage medium

The invention provides a document segmentation method and device based on a large language model, equipment and a storage medium, and relates to the technical field of text processing. The method comprises the steps of inputting a to-be-segmented target document into a pre-trained large language model, and executing the following operations through the large language model: performing text layout analysis on the target document, and identifying titles and all paragraphs of each level in the target document; for each paragraph, inserting an associated title related to the paragraph in all titles into an initial position of the paragraph to obtain a corresponding target paragraph; and sorting all the target paragraphs based on the semantic similarity among all the target paragraphs, and determining a segmentation result of the target document based on all the sorted target paragraphs. By the adoption of the technical scheme, when document segmentation is carried out, semantic loss in the document segmentation process can be effectively reduced, and therefore the document segmentation effect is improved.
Owner:CHINA LIFE ASSET MANAGEMENT CO LTD

Intelligent agent construction method and system based on large language model, equipment and medium

The invention provides an agent construction method and system based on a large language model, equipment and a medium, and belongs to the technical field of artificial intelligence. The method comprises the steps that a user interaction module is called to receive a first target document input by a user; calling a document analysis module to analyze the first target document to obtain a plurality of first chapter titles; calling an automatic mapping module to classify the plurality of first chapter titles based on a document classification model, and constructing a mapping relationship between the plurality of first chapter titles and a plurality of second chapter titles of a second target document based on a classification result; calling a large language model to respectively extract and summarize paragraph contents of the plurality of first chapter titles to obtain paragraph contents of a second chapter title; and calling a document output module to output the second target document according to a preset format. According to the agent construction method and system based on the large language model, the equipment and the medium provided by the invention, the accuracy of generating the review report by the agent can be improved.
Owner:ZHONGJIAO ROAD & BRIDGE (HEBEI) CO LTD

Ultra-long text generation method and device, equipment and storage medium

The invention discloses a super-long text generation method and device, equipment and a storage medium, and the method comprises the steps: responding to a document generation instruction, and generating an initial text outline of a target text through a large model; wherein the initial document outline comprises a plurality of chapter titles; based on the keyword information of the initial document outline, the text complexity of the target text is obtained through calculation, and the text complexity is associated with the subject breadth of keywords and the expected word number of the target text; if the text complexity of the target text is greater than a preset threshold value, determining the initial text outline as a target text outline; and on the basis of the target text outline, the text content corresponding to each chapter title is generated in parallel and spliced to obtain the target text, so that the generation efficiency, the text coherence and the overall text quality of the generated super-long text can be improved, and the generation cost of the super-long text is reduced.
Owner:SHENZHEN YUEHUA EXPRESS CO LTD

Text slicing and recall method and device, electronic equipment and storage medium

The invention relates to a text slicing and recall method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a to-be-processed document which comprises at least one original paragraph; for each original paragraph, semantic segmentation is carried out on the original paragraph to obtain at least one text block, and the word number of the text block is smaller than or equal to the maximum word number limit; according to a title structure of the to-be-processed document, a title path of each original paragraph is extracted, feature extraction is carried out on the title paths, title feature vectors are obtained, and the title paths are used for indicating titles of each level to which the original paragraphs belong; and for each text block, taking the title feature vector of the original paragraph to which the text block belongs as an index field of the text block to obtain a text slicing result, and using the index field to retrieve the text block. In this way, the semantically related text blocks are associated during retrieval, and the situation that related information is dispersed is reduced, so that the retrieval accuracy is improved, and the reliability of generating answers by the RAG system is improved.
Owner:BEIJING JINGTOU ZHUOYUE TECH DEV CO LTD

Automated Network Selection Management System

The headings and Abstract of the Disclosure provided herein are for convenience only and do not limit or interpret the scope or meaning of the embodiments. The various embodiments described above can be combined to provide further embodiments. Aspects of the embodiments can be modified, if necessary to employ concepts of the various patents, application and publications to provide yet further embodiments.Based on the teachings contained in this disclosure, it will be apparent to persons skilled in the relevant art(s) how to make and use embodiments of this disclosure using data processing devices, computer systems and / or computer architectures other than that shown in FIG. 4. In particular, embodiments can operate with software, hardware, and / or operating system implementations other than those described herein.Embodiments of the present disclosure have been described above with the aid of functional building blocks illustrating the implementation of specified functions and relationships thereof. The boundaries of these functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternate boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed.The foregoing description of the specific embodiments will so fully reveal the general nature of the disclosure that others can, by applying knowledge within the skill of the art, readily modify and / or adapt for various applications such specific embodiments, without undue experimentation, without departing from the general concept of the present disclosure. Therefore, such adaptations and modifications are intended to be within the meaning and range of equivalents of the disclosed embodiments, based on the teaching and guidance presented herein. It is to be understood that the phraseology or terminology herein is for the purpose of description and not of limitation, such that the terminology or phraseology of the present specification is to be interpreted by the skilled artisan in light of the teachings and guidance.The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments but should be defined only in accordance with the following claims and their equivalents.
Owner:SKYTELL AG

Retrieval enhancement generation method and device based on content hierarchical weighting and storage medium

The invention relates to the technical field of retrieval enhancement generation based on content hierarchical weighting, and discloses a retrieval enhancement generation method and device based on content hierarchical weighting and a storage medium, and the method comprises the following steps: constructing a document hierarchical tree, fully utilizing a father-son semantic relationship among titles, chapters and paragraphs in a document, and generating a retrieval enhancement result; according to the method, path weighting and node initial correlation scores are combined, the comprehensive score of each paragraph node is calculated, so that the correlation and importance of the selected target paragraph node are ensured, a fine-grained rearrangement technology ensures the logic continuity of information, the problem of evidence fragmentation is avoided, and finally, a reply text is generated through analysis of recalled data by a large model. The method has the advantages that the generated reply text is more accurate and effective, actual requirements of users can be met, retrieval result quality is improved, and user experience is enhanced.
Owner:SHENZHEN MINGXIN DIGITAL TECH CO LTD

Experimental commentary real-time generation method

The invention discloses an experimental explanation real-time generation method, which comprises the following steps that: a training database is constructed, a knowledge enhancement language model is trained, and the training database comprises an experimental title, a video, an audio, voice transcription text information and sensor time sequence data; the training process specifically comprises the following steps: splicing and fusing data in a database to obtain a structured input sequence, inputting the structured input sequence into a knowledge enhancement language model, and generating a corresponding basic operation instruction; the knowledge enhancement language model judges whether a knowledge module needs to be called for content supplementation or not, the content comprises scientific principles or operation specifications, and an experiment explanation text is output; and generating an experimental explanation speech corresponding to the experimental explanation text by adopting an end-to-end real-time speech generation mechanism based on semantic driving. The system has a real-time response capability, and can output voice content which is moderate in information amount, clear in structure and rich in guiding significance in real time according to sudden changes, dynamic events and the like in the experiment process.
Owner:SOUTH CHINA UNIV OF TECH

Method for generating offer information based on language model

The invention discloses an offer information generation method based on a language model, which relates to the field of natural language processing, and comprises the following steps of: retrieving news data from a multi-source corpus, forming an original data set, selecting field information of a title, an abstract, a text, a link, a keyword and a data source from the original data set, and storing the selected field information into a database; carrying out null removal and duplicate removal processing on the title and the text field information, and screening data sources to obtain a to-be-processed data set; on the basis of the to-be-processed data set, text vectorization processing is conducted on the title fields through a language model pre-trained by Transform, sentence vectors are obtained, and a sentence vector set is formed; according to the method, a high-quality to-be-processed data set is constructed through weighted Boolean retrieval and MinHash duplicate removal technologies, in the semantic analysis stage, sentence vectors are generated by adopting a RoBERTa model, unsupervised topic discovery is realized in combination with UMAP dimensionality reduction and HDBSCAN clustering, and the problem of high-dimensional text clustering is effectively solved.
Owner:AGRI INFORMATION INST OF CHINESE ACAD OF AGRI SCI

Index chain construction method based on RAG framework

The invention discloses an index chain construction method based on an RAG framework, which relates to the technical field of document processing, and comprises the following steps: receiving and analyzing different types of documents; uniformly converting the content into a Markdown format through an OCR (Optical Character Recognition) model, a layout recognition model and a table recognition model; segmenting the Markdown format document into a plurality of chunks, ensuring that a title and a text are logically coherent by utilizing a semantic analysis technology, and independently extracting a table to generate a natural language abstract; performing semantic analysis on each chunk by using a language model, extracting a high-dimensional semantic vector, processing an overlong text through a sliding window, and endowing a title vector with a higher weight to form a knowledge path; for chunk with less content, an adjacent block merging strategy is adopted to optimize a vectorization effect; and creating an index by adopting a cosine similarity method. According to the index chain construction method based on the RAG framework, the problem that in the prior art, the information processing accuracy of unstructured or semi-structured documents is low is solved, and the processing accuracy of the unstructured or semi-structured documents can be improved.
Owner:NAT SUPERCOMPUTING WUXI CENT

Role-aware passage theme event argument extraction method and device

The application provides a role-aware chapter theme event argument extraction method and device, and the method comprises the following steps: obtaining argument role information of a chapter theme event of an event type according to the event type; performing sentence segmentation and title extraction on a target article to obtain a sentence set and an event title; the argument role information, the event type, and the event title constitute event-related information; constructing an argument role-aware graph by using the event-related information and the sentence set, performing event-related sentence detection, and obtaining a chapter theme event-related sentence set; taking the chapter theme event-related sentence set as input, constructing a question for each argument role, predicting all candidate arguments in the chapter theme event-related sentence set, and screening out a target argument from the candidate arguments. The method improves the model effect while maintaining the flexibility of the model.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

System and Method for Automatic Data-Type Detection

A system and method utilizes masked language models in order to provide data-type detection, such as (but not limited to) prediction of columnar headings. Two masked language models are pre-trained on example columnar text. One model predicts missing data at the entity level (e.g., masked entity names that may be made up of whole words), while the other predicts missing data at the character level (e.g., masked individual characters). The table with missing column headings is fed into both models, and the output is contextual word embeddings and contextual character embeddings. These results are merged, and then fed into a neural network classifier to then predict the column names.
Owner:LIVERAMP

Electric power industry technology frontiers and personalized recommendation of articles of application thereof

The application provides a power industry science and technology front and application article personalized recommendation method and system, a storage medium and an electronic device, and relates to the field of personalized recommendation.In the application, the related field coefficient is used to screen sentences related to the personalized knowledge graph in the article to be recommended, and a new text after information filtering is generated; the title of the article to be recommended and the new text are used as the input of the BERT model, the title representation and the sentence representation are obtained respectively, the attention mechanism at the sentence level is introduced, the attention weight of the title representation to each sentence representation is obtained, the weighted sentence representation is obtained, and finally the classifier is input.The application uses the constructed personalized knowledge graph to screen sentences, preliminarily filters out irrelevant information, introduces the sentence-level attention model, filters the preliminarily screened sentences again, greatly reduces the interference of information irrelevant to the theme, and improves the performance of the classification model, so that the personalized recommendation result with high accuracy is quickly generated.
Owner:HEFEI UNIV OF TECH +1

Test case implementation method and device, terminal and storage medium

The invention relates to the technical field of testing, and discloses a test case implementation method and device, a terminal and a storage medium, and the method comprises the steps: receiving a test case title; wherein the test case title comprises a preset number of semantic fields, and each semantic field comprises a keyword and a parameter value corresponding to the keyword; analyzing the test case title to extract keywords and parameter values in each semantic field, and constructing a parameter dictionary based on an extraction result; and generating a configuration file for driving the test tool based on the parameter dictionary, and calling the test tool based on the configuration file to execute the test operation. The standardized test case title is automatically analyzed into the parameter dictionary and the test tool configuration file is generated, so that automatic conversion from the test intention to the execution task is realized, the development efficiency is remarkably improved on the premise of ensuring the test accuracy, and the maintenance cost is reduced.
Owner:SHENZHEN CITY TECHWIN SEMICONDUCTOR COMPANY LIMITED