Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

30 results about "Headword" patented technology

A headword, head word, lemma, or sometimes catchword, is the word under which a set of related dictionary or encyclopaedia entries appears. The headword is used to locate the entry, and dictates its alphabetical position. Depending on the size and nature of the dictionary or encyclopedia, the entry may include alternative meanings of the word, its etymology, pronunciation and inflections, compound words or phrases that contain the headword, and encyclopedic information about the concepts represented by the word.

Role-aware passage theme event argument extraction method and device

ActiveCN117194672BSemantic analysisKnowledge based modelsSentence segmentationEvent type
The application provides a role-aware chapter theme event argument extraction method and device, and the method comprises the following steps: obtaining argument role information of a chapter theme event of an event type according to the event type; performing sentence segmentation and title extraction on a target article to obtain a sentence set and an event title; the argument role information, the event type, and the event title constitute event-related information; constructing an argument role-aware graph by using the event-related information and the sentence set, performing event-related sentence detection, and obtaining a chapter theme event-related sentence set; taking the chapter theme event-related sentence set as input, constructing a question for each argument role, predicting all candidate arguments in the chapter theme event-related sentence set, and screening out a target argument from the candidate arguments. The method improves the model effect while maintaining the flexibility of the model.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

A training method of a source evaluation model, a source evaluation method, and related products

PendingCN122285891ARealize evaluationEvaluation resultMedicine
This application discloses a training method for a source evaluation model, a source evaluation method, and related products. Text sample sequences and frequency sample sequences are input into the evaluation model to be trained. The model encodes the text sample sequences and frequency sample sequences to obtain text sample vectors corresponding to the text sample sequences and frequency sample vectors corresponding to the frequency sample sequences. The evaluation model then performs prediction and evaluation processing on the text sample vectors and frequency sample vectors to obtain source evaluation prediction results. Based on the difference between the source evaluation result labels and the source evaluation prediction results, the parameters of the evaluation model are adjusted until the adjusted model meets the model training cutoff condition, and training ends to obtain the source evaluation model. Thus, this application can construct a source evaluation model based on the source sample name and the title, keywords, and publication frequency of the published sample text as features, thereby achieving the evaluation of the source.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

A scientific literature clustering method and system based on unsupervised keyword extraction

ActiveCN117453912BDegree of similarityHeadword
This invention relates to a scientific literature clustering method and system based on unsupervised keyword extraction. First, it effectively extracts keywords from scientific literature by comprehensively considering factors such as the occurrence of words in document abstracts and titles, the semantic similarity between words and the documents themselves, and the characteristics of domain keywords. Then, based on the characteristics of Chinese and English, this invention uses different embedding methods to cluster the extracted Chinese and English keywords, thereby achieving effective clustering of Chinese and English scientific literature. This invention considers the importance of words from multiple perspectives, comprehensively considering their occurrence in document abstracts and titles, and uses a method that automatically adjusts the preset keyword length based on domain characteristics to calculate keyword scores, thus incorporating more features of the words. This invention improves upon existing unsupervised keyword extraction algorithms.
Owner:SHANDONG UNIV

Enterprise intelligent knowledge base management system and method based on vector model

This invention discloses an enterprise intelligent knowledge base management system and method based on a vector model, relating to the fields of intelligent information retrieval and natural language processing. This invention parses unstructured documents in various enterprise formats, extracting semantic structure information such as text heading levels and paragraph boundaries, and segmenting them into a set of text fragments. The text fragments are then cleaned to obtain effective text fragments. These effective text fragments are integrated into semantic text blocks with the same topic. Sentence vectors are generated using a pre-trained semantic encoding model. Semantic segmentation points are located based on the semantic similarity of adjacent sentences to obtain standardized knowledge units. A dual-database collaborative query is constructed. After receiving user queries, structured pre-filtering is performed through a relational database, followed by semantic retrieval through a vector database. The intersection of the results from both databases is taken, and a comprehensive score is calculated based on multi-dimensional weights. After filtering, aggregation, and tracing, the results are output. If no matching results are found, a tiered fallback process is triggered, completing the entire retrieval loop.
Owner:XINFENG TECH (GUANGZHOU) CO LTD +2

Artificial intelligence-based title generation method, apparatus, device, and storage medium

The embodiment of the application belongs to the field of artificial intelligence, and relates to a title generation method based on artificial intelligence, comprising the following steps: acquiring a pre-acquired title data set; cleaning titles in the title data set based on multiple data cleaning rules to obtain target text data; using the target text data as training data to train a pre-trained language model based on an adversarial training strategy and a noise reduction strategy, to obtain a title generation model; acquiring a target text to be recognized; inputting the target text into the title generation model, and processing the target text by the title generation model to generate a corresponding target title. The application also provides a title generation device based on artificial intelligence, a computer device and a storage medium. In addition, the application also relates to blockchain technology, and the target title can be stored in the blockchain. The application uses the title generation model to quickly and accurately generate the target title corresponding to the target text to be recognized, and ensures the accuracy of the generated target title.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Ast-aware hybrid chunking corpus processing system for rag systems

The application discloses an AST-aware hybrid block processing system for RAG system and belongs to the technical field of natural language processing.The application comprises the following steps: parsing a Markdown document and constructing an abstract syntax tree; performing one-time segmentation according to the title level of the abstract syntax tree node type, performing secondary segmentation when the text block length exceeds a threshold value, and setting an overlapping area; cleaning the content of the text block; parsing YAML Frontmatter metadata and traceability information in the header of the Markdown document; connecting with an external server through an end-to-end cloud pipeline; and realizing multi-process parallel and streaming processing through a parallel processing architecture.The application can realize end-to-end automation and high-performance parallel processing under the premise of ensuring block accuracy and semantic integrity.
Owner:ZHEJIANG LAB

A large model-based text data governance method

The application relates to the technical field of electric digital data processing, in particular to a text data management method based on a large model. The method comprises the following steps: analyzing a target report text, and extracting structural information of the target report text; obtaining a probability that each title in a body part of the target report text refers to a table; extracting the content of each table in the table part, and generating an abstract description of each table by using a large model; matching each sentence corresponding to a candidate title in the body part with the abstract description and data of a candidate table; and if a certain sentence corresponding to the candidate title in the body part matches the abstract description and key data of the candidate table, inserting the corresponding position of the sentence into the identification of the candidate table. The application can automatically insert table identification in a report text, and improves the insertion efficiency.
Owner:山东省高级人民法院

A control unit for fine-tuning an intelligent / baselined model and a method thereof

The present application discloses a control unit for fine-tuning an intelligent / base model and a method thereof. The control unit 10 receives data from a plurality of sources 12 for training and developing an intelligent model 20 and then identifies each object present in each data point 11 of the received data and assigns a label to it by at least one expert model 15. The control unit 10 then derives metadata about each of the at least one expert model 15 and generates a plurality of titles using a large language model fuser 18. The control unit 10 sends the generated plurality of titles related to each data point 11 to a text encoder 16 for mapping at least one generated title to a corresponding data point 11. The control unit 10 then fine-tunes the intelligent model 20 using contrastive training techniques after mapping the data points 11 to the respective at least one generated title.
Owner:ROBERT BOSCH GMBH +1

WeChat public number tweet multi-modal question text inconsistency discrimination method and system based on causal inference

ActiveCN117892217BPublic accountData mining
The application relates to the field of artificial intelligence, and discloses a WeChat public account tweet multi-modal question-text inconsistency discrimination method and system based on causal inference, which comprises the following specific steps: acquiring training tweets and extracting multi-modal features; decoupling invariant factors and variable factors from the multi-modal features; based on a contrast learning strategy, decoupling the variable factors into causal factors reflecting deception tricks and writing styles of question-text inconsistency tweets under different scenarios and irrelevant factors containing false correlation bias; fusing the invariant factors and the causal factors; constructing a classifier; acquiring additional training tweets and performing data enhancement, training the classifier through the enhanced training tweets; identifying the titles and texts of the tweets through the trained classifier; and determining whether the tweets are question-text inconsistency. The application solves the problem that the prior art lacks modeling of false correlation bias, so that causal relationship information hidden in tweet features cannot be accurately explored, and has the characteristics of low demand for scarce labeled resources.
Owner:SUN YAT SEN UNIV

A method and system for constructing thinking chain data from a dissertation

The present application belongs to the field of artificial intelligence and natural language processing, and specifically relates to a method and system for constructing thinking chain data from a dissertation, aiming to solve the technical problems of high cost, low efficiency, unreliable quality and lack of traceable basis when constructing thinking chain data by artificial construction or directly generating thinking chain data by a large model. The present application includes: analyzing the title, table of contents and abstract of the dissertation to drive a large language model to condense key scientific problems; generating a core reasoning framework based on the problems and chapter directory, and extracting detailed argumentation content from the corresponding chapters step by step to integrate into a preliminary reasoning chain; using a long window large language model to compare and verify the reasoning chain and the full text of the dissertation and filter. The present application realizes the automatic and high-fidelity extraction of high-quality thinking chain data from the original dissertation, ensuring the logical rigor and traceability of the data.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Image capturing system, image capturing device, information processing server, image capturing method, information processing method, and computer program

An image capturing system configured to perform subject detection based on a neural network includes a teacher data input means, a network structure specification means, a dictionary generation means, and an image capturing device. The teacher data input means is configured to input teacher data for the subject detection. The network structure specification means is configured to specify a constraint on the network structure for the subject detection. The dictionary generation means is configured to generate dictionary data for the subject detection based on the teacher data and the constraint on the network structure. The image capturing device is configured to perform the subject detection based on the dictionary data generated by the dictionary generation means and perform predetermined image capturing control on a subject detected by performing the subject detection. The dictionary data includes, as header information, information regarding a count of teacher data used to generate the dictionary data.
Owner:CANON KK

Information processing system, terminal, information processing method, program, and data structure

PendingUS20260195330A1Search wordsInformation processing
A terminal and a computer system are provided, and in a case where a target search query of web content is received from a user, the terminal or the computer system outputs information for displaying a plurality of sub-search words / phrases used for search together with the target search query in another search query and / or a plurality of contained words / phrases mentioned on web pages up to a predetermined rank in search results of the target search query as a word / phrase list in a tree structure, and outputs information for displaying a combination of a heading tag assigned according to a hierarchy of each node of the tree structure and a word / phase of the node when a specific operation is executed by the user on a part or all of the displayed tree structure.
Owner:DATASCIENTIST

Spatial information processing system and method

A spatial information processing system according to an embodiment of the present invention may comprise the operations of: receiving an image and a query as inputs; extracting, from the image, a multi-dimensional visual feature representation including captions, object coordinates, and OCR text; converting the multi-dimensional visual feature representation into a linguistic representation understandable by a large-scale language model by using at least one lightweight artificial intelligence model; and receiving the converted linguistic representation and the query as inputs and generating a final response.
Owner:LG MANAGEMENT DEV INST CO LTD

A hierarchical search method, storage medium and computer device

This invention discloses a hierarchical retrieval method, storage medium, and computer device, belonging to the field of document processing technology. It includes: S100, obtaining the hierarchical paragraph parsing results of the input document; S200, retrieving the title paragraphs in the parsing results according to the user's query, generating a set of title paragraphs based on the retrieval results, determining whether the title paragraph set is empty, and if not, outputting a semantically complete paragraph; if so, proceeding to step S300; S300, retrieving the body text paragraphs in the parsing results according to the user's query, and outputting a semantically complete paragraph or summary based on the retrieval results. This invention, by combining the hierarchical parsing results of the document, can improve the completeness of text retrieval results in large-scale model retrieval and enhance document question answering, thereby improving the correctness and completeness of the generated answers in document question answering.
Owner:SHIP INFORMATION RES CENT (NO 714 RES INST OF CHINA STATE SHIPBUILDING CORP) +1

Large model-based knowledge base automated knowledge extraction and classification agent method

PendingCN122287825AData expansionLinguistic model
This invention discloses an automated knowledge extraction and classification agent method based on a large-scale model, relating to the field of artificial intelligence technology. The method includes knowledge base text training sample construction and annotation, structured preprocessing of knowledge base text, construction and training of a structure-first dual-channel knowledge extraction model, large language model-assisted data expansion, low-confidence difficult example correction, and online knowledge base text extraction. This invention employs a trigger-based correction strategy for low-confidence difficult example segments using a large language model, strictly limiting the scope of the large model call to structurally fragmented segments where model predictions are unreliable. It establishes a closed loop for correction result feedback and incremental fine-tuning, controlling computational overhead and response latency while ensuring overall extraction accuracy. A structural connection matrix based on document physical layout and hierarchical paths is constructed, explicitly encoding the non-linear positional relationships between headers and cells, and between titles and body text, as inter-word connection weights, enabling the model to directly utilize layout logic for semantic aggregation.
Owner:JILIN YOUYUN DIGITAL TECHNOLOGY CO LTD

Reviewer recommendation method for reasoning and structure dual-channel information enhancement

The invention relates to a reviewer recommendation method for reasoning and structure double-channel information enhancement, which effectively solves the problem of insufficient text information by designing a reasoning enhancement channel, deducing subject classification by using a large language model, generating methodology reasoning and revealing missing deep semantic features in titles and abstracts. A semantic similarity graph is constructed through a structure enhancement channel, a neighbor selection strategy is adopted, the papers are placed in a wider research background, and the range difference between the papers and professional knowledge of reviewers is effectively bridged. Information of the two channels is uniformly fused through a language model interpreter, and enhanced paper embedding is generated after fine adjustment on a subject classification task. According to the method, the professional knowledge portrait is constructed based on the historical manuscript review record of the reviewer, a graph neural network matching mechanism is introduced, and a cooperative signal is propagated through a two-part interaction graph, so that accurate matching of the thesis and the reviewer is efficiently realized.
Owner:TIANJIN UNIV

Multilingual trade document information extraction method and system based on artificial intelligence

PendingCN122290155Aprecise positioningImprove the success rate of extractionBorder lineEngineering
This invention relates to the field of document extraction technology, providing a method and system for extracting multilingual trade document information based on artificial intelligence. The method includes the following steps: translating table titles to obtain title content; identifying displayed and hidden border lines in the table to obtain the content of each cell; performing text matching between the extracted items and the content of each cell to determine the target cell with the highest matching degree; determining whether the target cell content and the extracted items are completely matched; if completely matched, determining the content of the right and bottom adjacent cells of the target cell as the content to be extracted; if not completely matched, reducing the extracted items according to the title content and then determining whether they are completely matched again; if not, breaking down the extracted items into words to determine the cells corresponding to each word, identifying the intersection area, and obtaining the content to be extracted. By adopting a hierarchical and progressive matching strategy, various matching situations are adaptively handled, resulting in better performance.
Owner:WUCHAN ZHONGDA INTERNATIONAL TRADING GROUP CO LTD

Intelligent generation method of keywords of Chinese and foreign language documents

The application discloses a kind of Chinese and foreign language literature keyword intelligent generation method, it is related to keyword extraction technical field.The present application includes the following steps: determining classification dimension and carrying out artificial marking in the Chinese and foreign language literature;Chinese and foreign language literature text data are cleaned, and sample data are obtained;With "keyword" as the class label of Chinese and foreign language literature, neural network with text classification as target is trained;TextRank algorithm and TF-IDF algorithm are used to extract text keyword;The extracted keyword is spliced behind title respectively, as data set;Keyword generation model is trained using data set pair;The output result of keyword generation model is analyzed, and accuracy is judged.The present application extracts Chinese and foreign language literature text keyword by improved TextRank algorithm and TF-IDF algorithm, inputs keyword generation model after keyword and title are fused and optimized, analyzes the output result of keyword generation model, judges accuracy, improves keyword generation method accuracy and coverage.
Owner:INST OF SCI & TECHN INFORMATION OF CHINA

System and method for creating a controllable output summary from text

ActiveUS12688222B2Data setEngineering
A system for creating a controllable output summary of text is disclosed. The system generates a set of summaries of text and for a first summary from among the set of summaries, executes a script that is configured to append the first summary to a set of summary-title pairs. The system generates a first title associated with the first summary in response to executing the script. The system compares the first title with the text. Based at least on the comparison, the system determines if the title indicates the context of the text. If it is determined that the title indicates the context of the text, the system generates a dataset of title-text pairs including the first title paired with the text. The system trains a target summarization algorithm with the generated dataset.
Owner:BANK OF AMERICA CORP

Providing generative artificial intelligence (AI) content based on existing in-page content in a workspace

PendingAU2026205406A1Patent Cooperation TreatyGenerative process
(12) INTERNATIONAL APPLICATION PUBLISHED UNDER THE PATENT COOPERATION TREATY (PCT) (19) World Intellectual Property Organization International Bureau (10) International Publication Number (43) International Publication Date WO 2025 / 096028 A1 08 May 2025 (08.05.2025) WIPOIPCT (51) International Patent Classification: (72) Inventors: LU, He; 2300 Harrison Street, San Francisco, G06F 40 / 10 (2020.01) G06N 20 / 00 (2019.01) California 94110 (US). SCALES, Jordan; 2300 Harrison G06F 40 / 166 (2020.01) Street, San Francisco, California 94110 (US). VARMA, At- ul; 2300 Harrison Street, San Francisco, California 94110 (21) International Application Number: (US). PCT / US2024 / 038250 (74) Agent: SIITONEN, Anni et al.; P.O. Box 1247, Patent Pro- (22) International Filing Date: curement, Seattle, Washington 98111-1247 (US). 16 July 2024 (16.07.2024) (81) Designated States (unless otherwise indicated, for every (25) Filing Language: English kind of national protection available): AE, AG, AL, AM, (26) Publication Language: English AO, AT, AU, AZ, BA, BB, BG, BH, BN, BR, BW, BY, BZ, CA, CH, CL, CN, CO, CR, CU, CV, CZ, DE, DJ, DK, DM, (30) Priority Data: DO, DZ, EC, EE, EG, ES, FI, GB, GD, GE, GH, GM, GT, 63 / 594,524 31 October 2023 (31.10.2023) US HN, HR, HU, ID, IL, IN, IQ, IR, IS, IT, JM, JO, JP, KE, KG, 18 / 408,429 09 January 2024 (09.01.2024) US KH, KN, KP, KR, KW, KZ, LA, LC, LK, LR, LS, LU, LY, 18 / 408,449 09 January 2024 (09.01.2024) US MA, MD, MG, MK, MN, MU, MW, MX, MY, MZ, NA, 18 / 408,454 09 January 2024 (09.01.2024) US NG, NI, NO, NZ, OM, PA, PE, PG, PH, PL, PT, QA, RO, 18 / 408,479 09 January 2024 (09.01.2024) US RS, RU, RW, SA, SC, SD, SE, SG, SK, SL, ST, SV, SY, TH, (71) Applicant: NOTION LABS, INC. [US / US]; 2300 Harri- TJ, TM, TN, TR, TT, TZ, UA, UG, US, UZ, VC, VN, WS, son Street, San Francisco, California 94110 (US). ZA, ZM, ZW. (54) Title: PROVIDING GENERATIVE ARTIFICIAL INTELLIGENCE (AI) CONTENT BASED ON EXISTING IN-PAGE CON- TENT IN A WORKSPACE 300 Compan. / / Al bio. Share is 305 AI Blocks 306 Al blocks are placed like ordinary Notion blocks, but contain as part of their metadata a prompt to a large language model. By default, the blocks display their intent (e.g. Summary this page) and a button to do SO. When the result of the large language model appears, the block updates to be a "container" of the AI output, with the ability to edit the prompt and re-run... + # "= Summarize this page with AI 302 Generate 304 309 308 Link The AI feature provides advantages such as displaying AI output at a specific position in the page and WO 2025 / 096028 A1 FIG. 3A (57) Abstract: A method for creating in-block content presented in a block on a page of a workspace. The block is configured to initiate a generative process to create in-block content of a particular type. The method includes determining a selection of in-page content based on a location of the block relative to the in-page content and the particular type of in-block content. The method can include causing a generative function to create generative content of the particular type based on the selection of the in-page content. The method can further include populating a block area to present the generative content. [Continued on next page] 20 26 20 54 06 08 J ul 2 02 6 0 8 J u l 2 0 2 6 W O 2 0 2 5 / 0 9 6 0 2 8 A 1 W I P O I P C T P C T / U S 2 0 2 4 / 0 3 8 2 5 0 2 0 2 6 2 0 5 4 0 6 U S M A , M D , M G , M K , M N , M U , M W , M X , M Y , M Z , N A , ( 5 4 ) T i t l e : P R O V I D I N G G E N E R A T I V E A R T I F I C I A L I N T E L L I G E N C E ( A I ) C O N T E N T B A S E D O N E X I S T I N G I N - P A G E C O N - 3 0 0 i s 3 0 5 3 0 6 + # A I 3 0 2 3 0 9 3 0 8 W O 2 0 2 5 / 0 9 6 0 2 8 A 1 F I G . 3 A
Owner:NOTION LABS INC

A method, apparatus, device and medium for structured processing of text

This application relates to the field of text structuring processing, and in particular to a method, apparatus, device, and medium for text structuring processing. In the process of text block structuring, fields in the previously generated structure tree are strategically masked, retaining only the true content of the last few valid nodes. This retained content is used to determine whether the current text block is a semantic continuation of the previous block, thus avoiding misjudgment of the title level. Based on the text instructions, the text blocks to be processed in this iteration, and the mask structure of the previous iteration, prompt words for the current iteration are constructed. This significantly compresses the context length input to the LLM in each iteration. Then, based on a large language model, the incremental structure tree for this iteration is output. Global hierarchical consistency is maintained through incremental expansion rather than reconstruction, avoiding root splitting. The nodes of the incremental structure tree are traversed, and the true content of the mask nodes of the incremental structure tree is retrieved and filled according to the content mapping table of the previous iteration, achieving accurate restoration.
Owner:ZHEJIANG TONGHUASHUN INTELLIGENT TECH CO LTD

A complex table question answering method based on tree structure, electronic equipment and medium

ActiveCN120670452BClearly understand hierarchical relationshipsUnderstand hierarchical relationshipsTable (information)Linguistic model
The application discloses a complex table question and answer method based on a tree structure, an electronic device and a medium, and comprises the following steps: in response to table header information of a complex table, a large language model outputs a table header tree structure; the large language model decomposes a question into a plurality of keywords, aligns the keywords with the table header tree structure, and obtains keyword-tree structures; in response to the keyword-tree structures, the large language model is guided to iteratively reason and search in the complex table by using a React-Style prompting method, and an answer is output. According to the application, the multi-level titles of the table are organized into a tree structure, so that the large language model can more clearly understand the hierarchical relationship of the table, and thus can more accurately locate relevant information when answering questions.
Owner:ZHEJIANG UNIV

Text processing model training method and apparatus, title generation method and apparatus, device, and medium

PendingEP4546206A4AlgorithmHeadword
Provided in the present disclosure are a text processing model training method and apparatus, a title generation method and apparatus, a device, and a medium. The text processing model training method comprises: acquiring a plurality of pieces of first title text, and inserting a mask segment into each piece of first title text multiple times, so as to obtain a plurality of pieces of second title text, into which the mask segment has been inserted, wherein the position where the mask segment is inserted into the first title text each time is different, and the number of characters of each piece of first title text is less than or equal to a preset number of characters; performing mask restoration processing on each piece of second title text on the basis of a trained mask restoration model, so as to obtain third title text after the first title text is expanded; and on the basis of the plurality of pieces of first title text and the plurality of pieces of third title text corresponding to the respective pieces of first title text, performing training to obtain a text processing model, which is used for performing title shortening processing, wherein during the training of the text processing model, the third title text is taken as input data, and the first title text is taken as output data.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Data verification methods, devices, equipment, and storage media based on OCR and NLP

This invention belongs to the field of artificial intelligence and provides a method, apparatus, device, and storage medium for document review based on OCR and NLP. The method includes: acquiring images of underwriting documents and image reference information; identifying face regions and character regions from the document images; when a document image is determined to be valid based on the proportion of the face region, identifying target characters and content characters using a preset OCR model, and determining a target dictionary based on the title characters of the target characters; inputting the target dictionary and content characters into a preset NLP model to identify the document review result. According to the technical solution of this embodiment, the validity of document images can be initially screened based on face proportion, eliminating the need for manual verification and reducing labor costs. Furthermore, after determining that the document image is valid, the OCR model and NLP model analyze the document review result, realizing automatic document image review and improving work efficiency.
Owner:CHINA PING AN LIFE INSURANCE CO LTD

A clickbait detection method using only title for prompt learning

ActiveCN117033639BDigital data information retrievalSemantic analysisStudies in Natural Language ProcessingPrediction score
The application discloses a click bait detection method for prompt learning only by using titles in the field of natural language processing research, comprising the following steps: 1. selecting a suitable pre-training language model as a backbone to construct label words and templates in prompt learning; 2. optimizing the label words in prompt learning through five optimization strategies, converting the classification task into a probability calculation problem of the category label words by using the expanded label words; 3. constructing the input text and the preset prompt template into a prompt text with mask as the input of the model, and performing click bait detection by using the optimized label words; 4. finally, mapping the predicted probability of each label word to the corresponding category to obtain the final prediction score of the label as the classification result; the application can obtain more accurate detection results by using fewer data through the five optimization strategies for screening prompt learning label words, greatly reduces the training cost of the model, and has high practicability.
Owner:YANGZHOU UNIV

Dynamic chapter generation for a communication session using an overall score and a chapter score

Methods and systems provide for providing dynamic chapter generation for a communication session. In one embodiment, the system connects to a communication session with a number of participants; receives a transcript of a conversation between the participants produced during the communication session, the transcript including timestamps for a number of utterances associated with speaking participants; segments the utterances into a plurality of contiguous topic segments; generates a title for each topic segment, the generating comprising: labeling one or more of the topic segments based on one or more labeling rules, extracting one or more top phrases from each topic segment, determining a ranking of the top phrases for each topic segment, and generating the title for each topic segment based on the top ranked phrase for the topic segment; and transmits, to one or more client devices, a list of the topic segments with generated titles for the communication session.
Owner:ZOOM COMMUNICATIONS INC

An english news element extraction method based on core word diffusion

The application provides a WHO element extraction method in an English news scene, which comprises the following steps: cleaning and preprocessing data of network news, combining an article title, a lead and an article theme into one text data, and analyzing a tree structure of the text data; extracting all nouns in all noun phrases from a subtree of the text data tree structure, screening core words, recalling a core word set, calculating a characteristic value of each core word, and extracting a first core word in a sorting result as a final core word; expanding the core word into a complete element text, traversing leaf nodes corresponding to the core word, and finding simple sentence nodes in a parent node direction, connecting all texts from the simple sentence nodes to the leaf nodes and then to intermediate links, and finally generating the element text; and the application realizes extraction of WHO elements on multiple news samples, has high accuracy, can improve a recall rate of core words, and can expand a recall set.
Owner:XIAMEN UNIV

Method and system for automatically generating reusable APIs based on code snippets

ActiveCN117892031BCode snippetEngineering
This invention discloses a method and system for automatically generating reusable APIs based on code snippets. Specifically, it includes: first, acquiring webpage data to be processed and extracting information contained in the webpage data, including the question title, several Q&A posts corresponding to the question title, and code snippets corresponding to each Q&A post; constructing prompt words for a large language model, including role-specific prompt words, thought chain reasoning prompt words, and few-shot context learning prompt words; inputting the constructed prompt words and each code snippet into the large language model, outputting an API method with a text description; and using regular expressions to clean the API method with the text description, removing the text description from the API method to obtain the API method corresponding to each code snippet. Compared with existing technologies, the method proposed in this invention can generate more accurate APIs.
Owner:SHANGHAI INST FOR ADVANCED STUDY OF ZHEJIANG UNIV