Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

303 results about "Headword" patented technology

A headword, head word, lemma, or sometimes catchword, is the word under which a set of related dictionary or encyclopaedia entries appears. The headword is used to locate the entry, and dictates its alphabetical position. Depending on the size and nature of the dictionary or encyclopedia, the entry may include alternative meanings of the word, its etymology, pronunciation and inflections, compound words or phrases that contain the headword, and encyclopedic information about the concepts represented by the word.

Science and technology public text intelligent classification and service method and device based on deep learning

The invention discloses a science and technology public text intelligent classification and service method and device based on deep learning. The method comprises the steps that multi-source science and technology public text data are acquired and preprocessed; extracting keywords by adopting a keyword extraction algorithm, and splicing the keywords with the public text title to form enhanced text features; constructing a multi-dimensional public text classification system and performing data annotation; performing feature extraction and fine adjustment by adopting a BERT pre-training model to obtain a classification model; automatically classifying the newly-added public texts and visually presenting the newly-added public texts; and generating a personalized recommendation result based on the user portrait and the public text feature index. The invention further relates to a technical scheme of multi-objective quality diversity optimization, heterogeneous resource allocation and fusion of the LPLC2 neural network and the BERT. The technical problems that a traditional method is limited in complex semantic understanding ability, single in classification dimension and lack of an integrated solution are solved, and the accuracy of science and technology public text classification and the intelligent level of service are improved.
Owner:GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS

Semantic enhancement adaptive partitioning method and system for natural resource large model questions and answers

The invention provides a semantic enhancement adaptive partitioning method and system for natural resource large model questions and answers, and aims to solve the problems of difficulty in term boundary recognition, damage to semantic integrity and the like. According to the method, three core technologies including theme perception coarse-grained paragraph division, self-adaptive sliding window theme hierarchy division and embedded perception context self-adaptive text segmentation are fused. The system firstly analyzes a natural resource long text structure, identifies titles and theme levels and aligns associated contents; paragraphs are extracted according to a theme perception strategy and are subdivided into sentence sets according to grammar rules; an improved sliding window mechanism is adopted to divide sentences into window sentence block groups. The method is characterized in that a dynamic aggregation threshold mechanism is introduced, the semantic association degree between adjacent sentence blocks is calculated through an embedded perception context semantic segmentation technology, whether the sentence blocks are combined or not is judged by combining a similarity distribution change trend and a dynamic adjustment threshold, self-adaptive delimitation of semantic boundaries is achieved, and text blocks which are clear in structure and coherent in semantics are generated.
Owner:HUBEI PROVINCIAL DEPT OF NATURAL RESOURCES INFORMATION CENT +1

Document content extraction method and system based on multimodal model collaboration, terminal and medium

The invention belongs to the technical field of document content extraction, and particularly discloses a document content extraction method and system based on multimodal model collaboration, a terminal and a medium. Comprising the following steps: identifying the type of an input to-be-processed document, and judging the document type; on the basis of the type identification result, calling a multi-modal model to analyze the document content, and outputting space coordinates, visual features and semantic features of document elements; generating a content sequence according with a reading habit through a semantic sequence reconstruction algorithm; paragraph boundary detection, paragraph recombination and semantic association modeling of charts and texts are completed based on the multilayer attention network and the graph neural network; grammar error correction, format optimization and title hierarchy generation are carried out by using a large language model and a hierarchical classification network; and converting the identification result into a structured output file. According to the method, the processing requirements of different types of documents can be considered, and high-precision analysis and efficient output are realized under the scenes of complex layouts, multiple languages and formula tables.
Owner:TUOSI (SHANDONG) INFORMATION TECHNOLOGY CO LTD

Highway intelligent operation and maintenance question-answering system based on large language model

The invention provides a highway intelligent operation and maintenance question-answering system based on a large language model, and belongs to the technical field of natural language processing. The system takes a large language model as a core reasoning engine and combines a domain knowledge base and an RAG technology to realize accurate question and answer of highway operation and maintenance; the method comprises the following steps: based on original knowledge data cutting, generating a title through a large language model, and customizing a knowledge base; receiving query, analyzing an intention by using a large language model, and matching to generate a function; the query is rewritten by using a large language model, dense and sparse vector query is generated, and a double-layer retrieval mechanism is formed; a two-step recall mode is utilized, coarse-grained recall is firstly carried out, then a recall result is subjected to fine-grained optimization through a screening mechanism, and a reasoning text is generated; and finally, inputting the query and reasoning text into the large language model, and generating an optimal answer through single-round and multi-round questions and answers. According to the method, the professionality and reliability of answers are enhanced, and the technical problem that answers are incomplete and inaccurate in an existing question and answer system is solved.
Owner:KUNMING UNIV OF SCI & TECH

Method and system for retrieving DOCX document content based on keywords

The invention belongs to the technical field of text processing, and particularly relates to a method and system for retrieving DOCX document content based on keywords, which comprises the following steps: analyzing an Office Open XML structure of a DOCX document, combining with multi-dimensional features such as style names, and utilizing a title classification score model to accurately distinguish a title and a text, so that a semantic hierarchical structure of the document is effectively reserved; and secondly, a multi-level semantic extension mechanism is introduced, and a Sension-BERT, a HowNet knowledge base and a Word2Vec model are fused, so that intelligent extension of synonyms and synonyms of keywords is realized, and the recall rate and semantic understanding ability of retrieval are remarkably improved. And in addition, a BM25 model is combined with paragraph length normalization and structure position weight to calculate a correlation score, so that retrieval results are sorted more accurately and reasonably. The construction of the reverse index is combined with the position coding and compression optimization strategy, and the retrieval efficiency and the storage performance are both considered.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Vehicle knowledge question-answering method based on GraphRAG method and computer equipment

The invention discloses a vehicle knowledge question-answering method based on a GraphRAG method and computer equipment, and the method comprises the steps: building a document tree structure based on the title hierarchy of a vehicle knowledge document, distributing a text identifier, and adding metadata information; converting the stored text units into text unit vectors by adopting a vector embedding model, and establishing a text vector index; calculating a similarity score between the query vector and the text unit vector, and determining a traditional RAG path text unit based on the similarity score; identifying entities and entity relationships of the text units according to the entity types and the relationship types by adopting a large language model so as to generate a vehicle knowledge graph; performing multi-level community division on the vehicle knowledge graph by adopting a Leiton community detection algorithm, and generating a structured community abstract according to each level of community; determining a GraphRAG path community abstract by adopting a large language model; and fusing the community abstract of the GraphRAG path with the text unit of the traditional RAG path and generating an answer.
Owner:FAW VOLKSWAGEN AUTOMOTIVE CO LTD

Intelligent writing method based on multi-agent cooperation

The invention discloses an intelligent writing method based on multi-agent cooperation, and the method comprises the following steps: receiving documents of various formats uploaded by a user, extracting text and image resources, and understanding an image through a multi-modal large language model to generate a description text; constructing a cue word template by using a large language model, processing three types of input files, extracting key information, and constructing a knowledge model; receiving user input construction enhanced cue words, and calling a large language model to generate an article outline; chapter contents are generated one by one according to a title segmentation outline, and a writing execution agent generates a text according to requirements and automatically illustrates; and the MarkDown format output is converted into Word with a basic format. The method has the beneficial effects that intelligent writing is completed through multi-step cooperation, information is extracted from multiple sources, a model is constructed, an outline and content are efficiently generated, a diagram is automatically illustrated, finally, the format is converted, and convenient, efficient and functional intelligent writing services are provided for users.
Owner:ZHEJIANG UNIV BINJIANG RES INST +2

Structured text analysis method and system based on title recognition and hierarchical abstract

The invention relates to a structured text analysis method and system based on title recognition and hierarchical abstracts, and the method comprises the steps: receiving a text, and obtaining sample data of a training language model; the vector database can perform vector dimension adjustment on the data of the stored text. The vector database takes vectors as basic storage units, and converts unstructured data into high-dimensional vectors through an embedding technology. And calling the data storage type information of the title text data set to determine the data storage type of the title text data set matched with the text storage mode. And obtaining a structure type of sample data of each text by utilizing a semantic segmentation script. And segmenting the document according to the identified chapter titles, and subdividing each chapter into a plurality of logic paragraphs to form structured text block units. The multi-level abstract generation model matched with the target model can fully mine level information and semantic association in the document. Therefore, the effect of analyzing the structured text by the intelligent question-answering system is improved.
Owner:HANGZHOU MEITENG TECH CO LTD

Long report generation method based on large model

The invention provides a long report generation method based on a large model, and belongs to the technical field of natural language processing and artificial intelligence, and the method specifically comprises the following steps: analyzing a report demand input by a user, extracting key information, and generating task description; automatically generating a report framework based on the large model and the domain knowledge base, wherein the report framework comprises chapter division, title generation and content summary; calling a large model to generate detailed contents in chapters, and introducing a context memory mechanism to ensure the continuity between the chapters; grammar check, logic check and style optimization are performed on the generated content, and the professionality and accuracy of the report are improved in combination with a domain term library; and finally, outputting the report content according to a format required by a user, and automatically generating auxiliary contents such as a directory, a chart and reference literature. The method has the advantages of high efficiency, high quality, flexibility, expandability and the like, and can be widely applied to long report generation tasks in the fields of medical treatment, law, finance and the like.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Retrieval enhancement generation method and system based on multivariate fusion

The embodiment of the invention provides a retrieval enhancement generation method and system based on multivariate fusion, and the method comprises the steps: firstly structuring a knowledge base document, generating a knowledge graph, and importing an ES to establish an index; inputting question sentences, carrying out word segmentation and other processing, querying instance nodes from the map, and outputting meeting conditions according to answer sentence patterns; if not, fragmenting and blocking the document according to a title level, obtaining candidate results through vector, ES and atlas retrieval, normalizing scores, merging and optimizing the scores of the candidate retrieval results according to a retrieval source, calculating comprehensive scores, and outputting the comprehensive scores in a descending order; and finally, carrying out semantic integrity aggregation on the combined candidate results, and inputting into a large language model to obtain a final answer. According to the method, the relevance of the retrieved content is greatly improved, the semantic integrity of the retrieved content is improved, the model magic view problem is reduced, and the question and answer accuracy and the answer quality of a knowledge base are improved.
Owner:BEIJING ZHITONG YUNLIAN TECH CO LTD

Water conservancy knowledge structured extraction and verification method and device

The invention provides a water conservancy knowledge structured extraction and verification method and device, and belongs to the technical field of artificial intelligence, and the method comprises the steps: carrying out the differential text processing of different formats of files, and generating an intermediate file; classifying the intermediate file into a regulation class or a non-regulation class based on a preset rule base; performing hierarchical title identification on the regulatory files to form entry knowledge blocks, and converting table contents into HTML (Hypertext Markup Language) knowledge blocks; performing semantic segmentation on the non-regulation file to generate knowledge blocks; performing knowledge block checking and filing, and marking an abnormal alarm block; converting the table knowledge blocks into natural language description by utilizing a large model; and positioning the context of the original text of the alarm knowledge block, and performing intelligent correction through a large model. According to the method, a traditional semantic analysis model and a large language model are creatively fused, a closed-loop process of preprocessing, extraction, verification and correction is formed, the problems of structured analysis and error correction of complex texts in the water conservancy field are solved, and the knowledge processing efficiency and accuracy are remarkably improved.
Owner:长江水利委员会网络与信息中心

Semantic understanding-based environmental impact report auxiliary auditing system

The invention relates to the technical field of semantic understanding, in particular to an environmental impact report auxiliary auditing system based on semantic understanding, which comprises a text cleaning module, a chapter segmentation module, a rule generation module, a content auditing module and a result output module. According to the method, the paragraph structure and the layout identification are analyzed in a unified mode, key paragraph types in environmental influence chapters are efficiently recognized, the text integration accuracy is improved by combining inter-paragraph similarity calculation and repetition rate screening, chapter boundary accurate positioning and affiliation adjustment are achieved on the basis of format feature comparison of serial numbers and title styles, and the text integration efficiency is improved. A multi-layer parameter alignment rule set is constructed to support comprehensive verification of standard numbers, monitoring frequencies and periodic elements, an exception labeling mechanism is combined to complete difference item logic judgment and consistency evaluation, an auditing basis with rule driving and data comparison capabilities is formed, chapter positioning precision, element extraction comprehensiveness and logic exception recognition efficiency are improved, and the method is suitable for popularization and application. And the pertinence and the automatic processing depth in the auditing process are enhanced.
Owner:SHANGHAI RUIDUN INFORMATION TECHNOLOGY CO LTD +1

Contract text key payment index extraction method based on natural language processing

The invention relates to the technical field of data processing, in particular to a contract text key payment index extraction method based on natural language processing, which comprises the following steps of: performing structured chapter division on a contract text, constructing a chapter hierarchical tree, and extracting a key payment index of the contract text according to a matching relationship between multi-level path nodes of the chapter hierarchical tree and chapter title keywords, setting a candidate path containing payment information; extracting an action entity of each contract text in the candidate path, and extracting a correlation index corresponding to the action entity according to a condition clause and a condition type of the action entity; extracting adjacent words based on the relevance index of the action entity, and constructing a word relation graph by analyzing the part-of-speech features of the adjacent words and the current action entity; iterating the word relation graph according to a preset association rule, and correcting each node in the word relation graph; outputting a key payment index corresponding to the contract text according to the difference part of the word relation graph before and after correction; accuracy and efficiency of payment index extraction are realized.
Owner:CHINA TIESIJU CIVIL ENGINEERING GROUP CO LTD +1

Distributed Hybrid Search for Language-Agnostic, Real-Time Information Retrieval

A computer-implemented method for performing searches in a document database is disclosed. The method comprises automatically detecting a line of business associated with a user, receiving a text query from the user, and generating a query embedding from the text query. The method further comprises scoring entries in a reverse index using a hybrid scoring function. The reverse index comprises titles, title embeddings, sentences, sentence embeddings, and entity tags corresponding to documents in the document database. The hybrid scoring function is used to generate a score based both on a keyword match score between the text query and the reverse index and on a cosine similarity score calculated from embeddings in the text query and in the reverse index. The method also comprises ranking scores for sentences in the document database, and displaying a sentence associated with a top score to the user.
Owner:DELL PROD LP

Knowledge base construction and retrieval method and system based on multi-source text in building field

The invention relates to the technical field of building information, and provides a knowledge base construction and retrieval method and system based on a multi-source text in the building field, and the method comprises the following steps: a knowledge base construction stage: constructing a multi-dimensional metadata feature vector for a multivariate text based on a standard classification index table; the method comprises the following steps of: converting and segmenting a document, splicing an end clause with all superior title texts by utilizing a context inheritance algorithm to form a text unit with complete semantics, and performing dynamic filtering based on an analyzed query intention and a metadata vector at a user retrieval stage; then, in the screening set, performing fusion calculation on semantic vector similarity, keyword matching degree and authority offset weight based on effectiveness attribute and implementation time, and performing mixed retrieval and reordering on the text units; and finally selecting a text unit according to a sorting result and inputting the text unit into the large language model to generate answers. According to the method, high-precision and high-compliance intelligent retrieval and question answering of building domain knowledge are realized.
Owner:SHANGHAI RESEARCH INSTITUTE OF BUILDING SCIENCES CO LTD

Complex table question and answer method based on tree structure, electronic equipment and medium

The invention discloses a complex table question-answering method based on a tree structure, electronic equipment and a medium, and the method comprises the following steps: responding to header information of a complex table, and outputting a header tree structure by a large language model; decomposing the question into a plurality of keywords through a large language model, and aligning the keywords with the header tree structure to obtain a keyword-tree structure; and in response to the keyword-tree structure, guiding the large language model to carry out iterative reasoning through a React-Style prompt method to retrieve from the complex table, and outputting an answer. According to the method, the multi-level titles of the table are organized into the tree structure, so that the large language model can more clearly understand the hierarchical relationship of the table, and related information can be more accurately positioned when questions are answered.
Owner:ZHEJIANG UNIV

Table text boundary adaptive information extraction method and system

The invention provides a form text boundary adaptive information extraction method and system in the technical field of computer information processing. The method comprises the following steps: S1, identifying a title frame from a form image; s2, calculating the aspect ratio of the title frame to determine the title arrangement direction of the table titles; s3, detecting a text detection box of the table image, and performing classification prediction of a text arrangement direction on the text detection box by combining a text direction model with a title arrangement direction; s4, controlling affine transformation of the table image through the text arrangement direction and the text detection box to obtain a corrected image; s5, identifying a table title in the corrected image, and matching a table type from the table type configuration file based on the table title; s6, extracting keywords from the corrected image based on the table type; and S7, processing the keyword to extract table information. The method has the advantages that the adaptability, accuracy and efficiency of table information extraction are greatly improved.
Owner:FUJIAN NEWLAND SOFTWARE ENGINEERING CO LTD

System and method for creating a controllable output summary from text

A system for creating a controllable output summary of text is disclosed. The system generates a set of summaries of text and for a first summary from among the set of summaries, executes a script that is configured to append the first summary to a set of summary-title pairs. The system generates a first title associated with the first summary in response to executing the script. The system compares the first title with the text. Based at least on the comparison, the system determines if the title indicates the context of the text. If it is determined that the title indicates the context of the text, the system generates a dataset of title-text pairs including the first title paired with the text. The system trains a target summarization algorithm with the generated dataset.
Owner:BANK OF AMERICA CORP

Knowledge base context awareness and traceability enhanced intelligent retrieval and question-answering system

The invention provides an intelligent retrieval and question-answering system for knowledge base context awareness and traceability enhancement, and the system obtains an analysis result, an optical character recognition result and a semantic analysis result through the analysis of PDF physical layout, Word / Excel paragraphs, titles, tables and style attributes thereof, and the optical character recognition. Structured knowledge blocks, positions and levels of path images and vector representation are generated and stored in a vector database, and related results of user query requests are extracted by executing a mixed retrieval algorithm with keyword retrieval and vector semantic retrieval. Performing intelligent reordering by considering semantic similarity, keyword matching, knowledge block type weight, source knowledge base weight and page position weight to obtain a candidate knowledge block list, constructing cue words according to an intelligent reordering result, and calling an external large language model to generate answers; source labels in answers are managed to be associated with metadata of corresponding numbers in a candidate knowledge block list, and the problem that text blocks and traceability are not accurate is solved.
Owner:CHINA HAISUM ENG

Data processing method and device, equipment and medium

The embodiment of the invention provides a data processing method and device, equipment and a medium, and the method comprises the steps: carrying out the content understanding of media data, obtaining a first content summary text, and generating M first content titles of the first content summary text through a first language model; obtaining interaction feedback data of each first content title in the M first content titles, and performing quality evaluation on each first content title according to the interaction feedback data of each first content title to obtain a quality evaluation value of each first content title; dividing the M first content titles into positive feedback titles and negative feedback titles according to the quality evaluation value of each first content title; performing parameter optimization on the first language model according to the first content abstract text, the positive feedback title and the negative feedback title to obtain a second language model; the second language model is used for optimizing a negative feedback title of the media data. By adopting the method, the title generation quality of the media data can be improved.
Owner:TENCENT TECH (BEIJING) CO LTD

Duplication check system and method for paper generated by artificial intelligence

A duplication check system and method for paper generated by artificial intelligence includes steps: S1: the user uploading the academic paper to be detected to a system, and the system automatically extracting the title, the abstract, and the headline of each paragraph of the paper; S2: fusing the title, the abstract, and the headline of each paragraph of the paper with the contextual information of the paper and extracting theme features; S3: after the different themes of the paper are extracted, repeatedly using similar tones for each theme in all different AI tools to propose integrate text requirements, searching each theme for times of the number of repetitions of integration in each AI tool until no new content is obtained, matching all the obtained texts with the paper to be duplication checked, based on natural language understanding, for duplication check, and marking the matching repeated parts and indicating the sources.
Owner:WU JIANG

Cross-modal fusion lightweight defect detection method based on knowledge distillation

The invention belongs to the technical field of digital image processing, and particularly relates to a knowledge distillation-based cross-modal fusion lightweight defect detection method, which comprises the following steps of S10, cross-modal fusion distillation; through a bidirectional vision-language alignment mechanism, the frozen multi-modal knowledge of a teacher model vision-language basic model VLM is migrated to a lightweight student model, and the dual-path fusion module comprises text condition region representation injected with semantic context and region anchoring semantic embedding fused with spatial vision clues; step S20, cross-header word-region alignment is carried out; embedding the fusion visual features generated by the two-way fusion module and the enhanced text to generate cross-head prediction so as to simulate the semantic-space association capability of a teacher model; s30, knowledge distillation loss is fused; according to the method, the multi-modal basic model is fused and distilled into the lightweight single-modal detection model, the detection performance in a defect detection scene can be improved, and compared with the basic model, the reasoning speed is greatly improved, and the parameter quantity is reduced.
Owner:CENT SOUTH UNIV

Jewelry value evaluation knowledge base construction method based on large language model

The invention discloses a jewelry value evaluation knowledge base construction method based on a large language model. The method comprises the following steps: S1, collecting and preprocessing jewelry value evaluation text data; s2, keeping title enhancement of the hierarchical information; s3, standardizing table data; s4, converting the text including the table into a Markdown format; s5, extracting jewelry value evaluation knowledge based on single sample learning by using a large language model; s6, processing and storing a result; according to the method, the title is converted into the Markdown format, the absolute path of each title in the text is reserved, and meanwhile, the cue word for guiding the large language model to extract information in a structured manner is constructed for the extraction task, so that structured extraction and tree structure organization of jewelry value evaluation knowledge are realized; the problem that a hierarchical structure is difficult to construct in a knowledge base due to the fact that a current text information extraction result based on a large language model lacks all levels of title link relations is solved.
Owner:INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS

Sparse N-gram modeling for patient-entity relation extraction

Methods, systems, and software are provided for determining a relationship between a subject and a health entity. An electronic health record (EHR) for the subject is split into sections by detecting delineating section headers, and sections are subdivided into text spans. Text spans are filtered by language pattern recognition into a set of text spans having an expression related to the health entity. The natural language context of the expression in each text span in the set is evaluated to obtain a corresponding scoring representation. Scoring representations are inputted into a model comprising a plurality of parameters. The model outputs, for each text span in the set, at least a prediction that the text span is associated with the health entity. Models for determining relationships between subjects and health entities and methods for training models to determine relationships between subjects and health entities are also provided.
Owner:TEMPUS AI INC

Book-based subject knowledge system construction method and device

The invention discloses a book-based subject knowledge system construction method and device, and the method comprises the steps: representing a directory as a directed tree, and dividing the directed tree into a skeleton module and a plurality of chapter modules; carrying out the association and expansion of a title through a large language model, and generating an abstract; knowledge concepts, subject terms and structural relations in titles in all the modules are extracted and subjected to double correction to obtain a plurality of knowledge system trees, and then a mother tree obtained after the skeleton module is processed and a sub-tree obtained by the chapter module are spliced; and finally, giving a plurality of knowledge trees obtained according to a plurality of groups of book directories in the same subject field, and combining the knowledge trees into a unified knowledge system tree based on a finite-state machine principle. According to the method, the knowledge system is constructed and decoupled into an extraction stage and a fusion stage based on LLM, the problem of a pain point of drifting between knowledge concepts and attributes when a simple chapter title is faced is solved, and strict control on semantic consistency and hierarchical rationality in a path merging process is ensured based on a state machine.
Owner:ZHEJIANG UNIV

Semantic recall model training method and device and storage medium thereof

The invention provides a semantic recall model training method and device and a computer storage medium. The method comprises the following steps: acquiring a training sample set, wherein the training sample set comprises a plurality of sample pairs and sample labels thereof; for each sample pair in the training sample set, respectively inputting the query text and the title text into a first semantic encoder and a second semantic encoder of a semantic recall model to obtain a first semantic vector corresponding to the query text and a second semantic vector corresponding to the title text; at least inputting a second semantic vector corresponding to the title text into a semantic decoder of a semantic recall model to obtain a predicted event text; extracting a third semantic vector of the prediction event text; determining a first loss based on the first semantic vector and the third semantic vector corresponding to each sample pair in the training sample set and the sample label; and iteratively updating the parameters of the semantic recall model until a preset condition is met. According to the method, the timeliness semantic recall effect can be effectively improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method and device for intelligent semantic error correction and business term optimization of foreign trade letter electricity

The invention relates to the technical field of natural language processing, in particular to a foreign trade letter intelligent semantic error correction and business term optimization method and device, and the method comprises the steps: obtaining a target foreign trade letter, and constructing a target corpus; performing Chinese word segmentation and part-of-speech tagging on the target foreign trade letter; performing term optimization based on the word segmentation result and the knowledge graph; identifying the letter title by using a conditional random field model, and converting the letter title into structured data; a Bi-LSTM-CRF model is adopted to carry out risk point detection, including Bi-LSTM coding, feature engineering and Max-pooling technologies, a part-of-speech sequence is obtained through a softmax function and a Viterbi path, and sequence labeling is carried out to obtain a risk point detection result; and finally, performing Chinese error correction based on the word segmentation result after part-of-speech tagging and the knowledge graph. The recognition and correction accuracy of foreign trade terminologies is improved, and communication obstacles caused by nonstandard use of the terminologies are effectively reduced.
Owner:GUANGDONG VOCATIONAL COLLEGE OF SCI & TRADE

A text analysis method, system, device and medium based on text structure

The present invention relates to a text analysis method, system, device and medium based on text structure, which comprises the following steps: parsing the acquired text to be analyzed to obtain its text structure; respectively performing machine reading on each text structure of the text to be analyzed to obtain the embedding vectors corresponding to the respective text structures; fusing the obtained embedding vectors to obtain a fused article embedding vector; and obtaining a text analysis result based on the fused article embedding vector. The present invention takes into account the important significance of the article structure for machine understanding, and parses according to the structure of abstract - paragraph {paragraph title - paragraph content}, enabling the model to have the ability of reading by sub - structure. Therefore, the present invention can be widely applied to the field of text analysis.
Owner:RENMIN UNIVERSITY OF CHINA

Method and equipment for extracting bill of material and title information in CAD (Computer Aided Design) and medium

The invention discloses a method and device for extracting a bill of material and title information in CAD and a medium, and the method comprises the steps: extracting an INSERT insertion object and an ATTRIB attribute object from a drawing object list of a CAD file, and taking the INSERT insertion object and the ATTRIB attribute object as filtering data; the BOM and the title information are extracted from the filtered data through a preset rule matching algorithm, and if extraction succeeds, the BOM and the title information are output; and if the extraction fails, extracting the BOM and the title information from the filtered data according to a preset dictionary matching algorithm, and outputting the BOM and the title information. According to the method, a dictionary matching and rule matching algorithm in NLP (Natural Language Processing) and text analysis is fully utilized to perfectly solve the pain point of manually inputting the BOM in the use process of an ERP (Enterprise Resource Management) system.
Owner:NANJING LETSTECH CO LTD

Document segmentation method and device based on large language model, equipment and storage medium

The invention provides a document segmentation method and device based on a large language model, equipment and a storage medium, and relates to the technical field of text processing. The method comprises the steps of inputting a to-be-segmented target document into a pre-trained large language model, and executing the following operations through the large language model: performing text layout analysis on the target document, and identifying titles and all paragraphs of each level in the target document; for each paragraph, inserting an associated title related to the paragraph in all titles into an initial position of the paragraph to obtain a corresponding target paragraph; and sorting all the target paragraphs based on the semantic similarity among all the target paragraphs, and determining a segmentation result of the target document based on all the sorted target paragraphs. By the adoption of the technical scheme, when document segmentation is carried out, semantic loss in the document segmentation process can be effectively reduced, and therefore the document segmentation effect is improved.
Owner:CHINA LIFE ASSET MANAGEMENT CO LTD