Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

81 results about "Formatted text" patented technology

Formatted text, styled text, or rich text, as opposed to plain text, has styling information beyond the minimum of semantic elements: colours, styles (boldface, italic), sizes, and special features in HTML (such as hyperlinks).

Engineering document index consistency proofreading method and system based on multi-modal large model

The invention relates to an engineering document index consistency proofreading method and system based on a multi-modal large model, and the method comprises the steps: Q1. OCR detection and recognition: carrying out the optical character recognition and format analysis of a source document, converting an uploaded PDF document into a processable text message in a Markdown format, and carrying out the format discrimination of a table, a formula and a plain text; and Q2, table and formula processing: adopting a hierarchical processing strategy, intelligently selecting an optimal processing mode according to the complexity of the table, and converting table information into a descriptive long text through a language large model and cue words. According to the method, accurate, reliable and efficient document index checking service can be provided for a user, the quality and efficiency of professional document processing are remarkably improved, the efficiency and quality of knowledge graph construction are remarkably improved, a knowledge verification system capable of being evolved continuously is established, and the method is suitable for popularization and application. And a reliable technical support is provided for knowledge management and professional decision-making in a complex field.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

Vertical field data construction method based on large model

The invention provides a vertical field data construction method based on a large model, which belongs to the technical field of data processing and artificial intelligence, and comprises the following steps: converting a vertical field source document into an intermediate format text, and segmenting the intermediate format text into a plurality of text blocks; inputting the text blocks into a pre-trained generative language model, guiding the pre-trained generative language model according to pre-designed cue words to generate a plurality of candidate questions according to the content of the text blocks, and performing preliminary screening and fine screening on each candidate question to obtain a question set; pre-defining a mode of a knowledge graph according to field characteristics of the vertical field, processing all text blocks based on an information extraction model, and constructing a field knowledge graph; and performing local context retrieval on each final question in the question set based on the text block of the question source, performing global knowledge retrieval based on the domain knowledge graph, and generating a final answer and a final thinking chain. The method is suitable for different vertical fields, the data quality can be effectively improved, and the problem generation accuracy is guaranteed.
Owner:PEKING UNIV

Text and table double-track collaborative data extraction method and system oriented to alloy field

The invention discloses an alloy field-oriented text and table double-track collaborative data extraction method and system. The method comprises the following steps of: obtaining an alloy related document DOI; taking the DOI as an index to obtain a full-text literature and storing the full-text literature in a JSON format; performing document screening based on an LDA topic distribution model and establishing a target topic document library; preprocessing the unstructured table, and obtaining structured table data by using a large language model; performing automatic text segment division on the JSON format text through rule matching analysis; and sorting the text segments by using a large language model to generate semi-structured text data, fusing and normalizing the text and table data, and outputting a component-process-organization-performance-strengthening mechanism alloy data chain. According to the method, through a double-track collaborative data extraction mechanism, the problems of multi-source heterogeneity, format fragments, information islands, semantic layering and the like of material subject data are effectively solved, structured treatment of literature data in the alloy field is achieved, and a unified technical path is provided for constructing a high-quality alloy database.
Owner:CENT SOUTH UNIV

Plant protection field literature information batch structured extraction method based on AI large model

The invention discloses a plant protection field literature information batch structured extraction method based on an AI large model, and belongs to the technical field of computer data processing and artificial intelligence application. According to the method, a PDF format literature is converted into a Markdown format text through a PDF literature self-adaptive preprocessing module and a Markdown conversion module; then, a double-strategy self-adaptive AI extraction module is adopted, different modes are adopted for processing according to a text length threshold value, and JSON data are extracted in combination with a JSON format data restoration strategy; then, the CPU intensive conversion task and the I / O intensive AI analysis task are processed in parallel through a two-stage parallel scheduling module, and a structured database file is generated through aggregation according to a predefined mapping rule through a multi-dimensional aggregation export module. According to the method, end-to-end automation from literature acquisition to structured data output is realized, and the problem of low extraction efficiency of manual literature reading information is solved.
Owner:SANYA INSTITUTE OF NANJING AGRICULTURAL UNIVERSITY

Method, device and equipment for constructing dynamic expansion word bank and medium

The invention relates to the technical field of passenger service, and discloses a construction method and device of a dynamic extension word bank, equipment and a medium. The method comprises the following steps: acquiring multi-source heterogeneous original knowledge data related to civil aviation passenger service, and converting the multi-source heterogeneous original knowledge data into a standard format text; processing the unstructured data in the standard format text based on the adaptive word segmentation model and the stop word list to obtain a plurality of candidate words; performing multi-dimensional corpus feature quantitative analysis on each candidate word to determine candidate keywords; and performing similarity calculation based on the entity word segmentation and the candidate keywords, determining effective candidate keywords, and adding the effective candidate keywords into the dynamic expansion word bank. By means of the method and device, the technical problems that in the prior art, a word bank used by a civil aviation service system is usually based on a static vocabulary or depends on manual compiling and updating, the static word bank cannot be updated in real time, the characteristics of the civil aviation field cannot be flexibly handled, and the maintenance cost is high due to manual updating and maintenance are solved.
Owner:TRAVELSKY TECHNOLOGY LIMITED

Intelligent cue word generation evaluation method and system based on multi-component collaboration

The invention discloses an intelligent cue word generation evaluation method and system based on multi-component collaboration, and relates to the technical field of artificial intelligence and natural language processing. The method comprises the following steps: acquiring a multi-format text, converting the multi-format text into an input-output pair, and generating a plurality of candidate cue words; after determining a communication mode according to the text form, standardizing cue words according to the structured template of each large language model; calling each platform to generate answer analysis and unify formats, and extracting target contents by utilizing regularization; and carrying out multi-mode evaluation on the candidate cue words, selecting an optimal cue word, and outputting a structured evaluation result and an optimization suggestion based on multiple indexes. According to the method, full-process automatic processing of prompt word generation, optimization and evaluation is achieved, various model platforms, data sets, algorithms and evaluation indexes are supported, and good universality and expansibility are achieved.
Owner:AIR FORCE UNIV PLA

Vertical field objective test data set construction method, system, equipment and medium

The invention discloses a vertical field objective test data set construction method, system and device and a medium, and relates to the technical field of natural language processing, the construction method comprises the following steps: obtaining text data, and carrying out knowledge extraction and arrangement on the text data to obtain formatted text information; taking the formatted text information as input, and performing question generation, correct answer generation, interference option generation and analysis generation through a pre-established large model to obtain test data; performing quality detection on the obtained test data, and screening out the test data meeting the quality detection requirements; and performing self-inspection on the test data meeting the quality detection requirements, and retaining the test data passing the self-inspection to obtain a vertical field objective test data set. According to the method, the vertical domain test questions can be automatically generated on a large scale, the dependency degree of testers on professional domain knowledge is reduced, meanwhile, the labor cost is saved, the problem difficulty is controllable, and the professionality and accuracy of test question generation are improved.
Owner:CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +2

Systems And Methods For Generative Language Model Database System Communication Channel Integration

A computing services environment may include a database system storing a plurality of database records for client organizations accessing computing services including a conversational chat assistant accessible via various communication channels. The computing services environment may also include a communication interface configured to receive an input message from a client machine via a communication channel, a generative language model interface providing access to one or more generative language models, and an orchestration and planning service. The orchestration and planning service may be configured to analyze the input message to determine a novel text passage via a generative language model, to determine novel text formatting information based on designated text formatting configuration information specifying one or more parameters for formatting text generated for transmission via the communication channel, and to transmit the novel text passage and the novel text formatting information to the client machine via the communication interface.
Owner:SALESFORCE INC

Industrial internet network intrusion detection method based on pre-trained large language model

The invention provides an industrial internet network intrusion detection method based on a pre-trained large language model, and the method comprises the steps: carrying out the data preprocessing of original network flow data through protocol self-adaptive flow aggregation and session segmentation, feature screening and feature coding; constructing a stream format text data set used for training a generation model GPT-2 and a packet level classification text-label data set used for training a classification model DistilBERT; then fine tuning training is carried out on the generation model GPT-2 and the classification model DistilBERT; then calling a trained generation model GPT-2 to generate a traffic sequence, obtaining a prediction data packet, calling a trained classification model DistilBERT, performing anomaly judgment on sequence data in the prediction data packet one by one, and outputting a classification result of anomaly judgment; the active intrusion detection of first generation and then discrimination is realized. According to the method, prediction can be made before real attack traffic arrives, and network attacks are prevented.
Owner:EAST CHINA JIAOTONG UNIVERSITY

Intelligent extraction method and system for input text containing mathematical formula

The present invention belongs to the technical field of text processing, and relates to an intelligent extraction method and system for an input text containing a mathematical formula. The method comprises: 1) performing format determination, conversion and preprocessing on an input text; 2) performing angle correction on the preprocessed image-formatted text; 3) performing formula detection; 4) performing layout analysis; 5) for inline formulas, according to a formula detection box, determining whether a corrected OCR detection box contains an inline formula, and splitting the OCR detection box containing the inline formula, so as to obtain a text-only OCR detection box; 6) performing formula recognition to obtain a formula recognition result; 7) performing text recognition to obtain a text recognition result; and 8) on the basis of a layout analysis box and a layout category thereof, performing same-line detection box determination and merging on the formula recognition result and the text recognition result, so as to obtain an extraction result for the input text. The present invention can effectively improve the extraction efficiency and accuracy for input texts containing mathematical formulas.
Owner:BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD

Aviation text content cleaning and labeling method, system and equipment and medium

PendingCN121543549ANatural language analysisBiological modelsDuplicate contentAviation
The invention relates to the technical field of aeronautical text data processing, and discloses an aeronautical text content cleaning and labeling method, system, device and medium wherein the method comprises: noise filtering: identifying and removing noise in an aeronautical text in combination with static cleaning and a general large model; format standardization: converting the aviation text after noise removal into a standardized format text; duplicate removal and error correction: detecting duplicate contents based on a hash algorithm, and correcting spelling errors and grammar errors based on a general large model to obtain an aviation text subjected to duplicate removal and error correction; entity identification: key entities are extracted based on the general large model, and the extracted key entities are labeled; active learning: screening high-value samples in the marked key entities based on an uncertainty query strategy; and dynamic optimization: performing verification and iterative optimization on the marking result in combination with the aviation knowledge base. According to the method, the automation level, the labeling accuracy and the system self-adaptive capability of aviation text processing can be remarkably improved.
Owner:四川腾盾科技有限公司 +1

Style and style aligned DOCX format document translation system

The invention discloses a style and style aligned DOCX format document translation system, which comprises the following steps of: analyzing a DOCX format document through a document analysis tool to obtain an XML format text; each Paragraph object style label in the XML format text is simplified; the simplified XML format text is sent into a large language model to obtain a translation XML format text; carrying out XML label integrity verification on each XML format text of the translation; if the verification result meets the integrity requirement, original style tags are restored according to the corresponding sequence of the text in the XML format of the translation; and generating the reduced text in the XML format of the translation into a document in the DOCX format of the translation through a document analysis tool. According to the method, the problem of balance between format fidelity and translation quality in professional document translation is effectively solved, the pattern alignment accuracy and translation efficiency are remarkably improved, and the method is particularly suitable for professional document translation scenes, such as legal contracts and technical manuals, in which formats need to be strictly reserved.
Owner:XIAONIU FANYI

Computationally efficient artifact tagging for document management

Image pages are generated from a document. Assembling the image pages generates a collage. Two-dimensional text and bounding boxes are extracted from the image pages. A structure verbalizer spatially formats the two-dimensional text in one-dimension with spatial information to generate spatial-formatted text. The spatial-formatted text is concatenated to generate a text extraction. A multimodal embedding model is applied to the collage and the text extraction to generate a target artifact vector. The target artifact vector is compared against a set of preexisting artifact vectors to identify a corresponding artifact vector associated with a corresponding document having a corresponding metadata tag. A distance is determined between the corresponding artifact vector and the target artifact vector. Responsive to the distance being within a threshold distance additional steps are performed, including performing both applying the corresponding metadata tag to the document to generate a modified document and outputting the modified document.
Owner:INTUIT INC

Contract information extraction method and device, equipment and medium

The embodiment of the invention relates to a contract information extraction method and device, equipment and a medium. The method comprises the steps of obtaining a to-be-processed clause text in a to-be-processed contract text; responding to the to-be-processed clause text as a target clause text, obtaining a prompt word template corresponding to the to-be-processed clause text, and setting a corresponding prompt identifier for the prompt word template corresponding to the to-be-processed clause text to obtain a target prompt word template; and performing information extraction on the to-be-processed term text based on the target cue word template through a preset large model to obtain target format text information so as to determine contract information corresponding to the to-be-processed contract text. Therefore, the key information of each clause text can be quickly and accurately extracted through the large model, the accurate information of the whole contract text can be obtained, the prompt identifier is added for the specific clause to improve the recognition accuracy, and the contract information recognition efficiency and effect are further improved; therefore, the contract-related business processing efficiency is improved, and the business error risk is reduced.
Owner:KE COM (BEIJING) TECHNOLOGY CO LTD

Adaptive text rendering system for digital content editing applications

PendingUS20260187339A1Digital contentUser input
A system for adaptive text rendering includes an editor, a context analyzer, an adaptive text generator, and a feedback and learning engine. The editor renders a text box on a digital canvas and receives user input via the text box. The context analyzer extracts contextual information from digital content associated with the digital canvas. The adaptive text generator uses the user input, text box dimensions, and contextual information to direct an artificial intelligence (AI) model to generate text. It then formats the generated text to fit within the text box based on one or more fit criteria. The feedback and learning engine captures and evaluates user feedback related to the formatted text and generates tuning adjustments to refine the adaptive text generator or the AI model.
Owner:WIX COM

Data query method, device and storage medium

This application discloses a data query method, device, and storage medium, relating to the field of Internet technology. This method addresses the existing problem of being unable to perform text queries based on query statements. The method includes obtaining a query statement and structured text data; parsing the query statement into an abstract syntax tree; attaching the structured text data as a parameter to the abstract syntax tree; compiling the abstract syntax tree to generate a high-order function; executing the high-order function to generate a row-level data access handle; using the row-level data access handle to access the structured text data, obtain query results, and return tabular data of the query results. This method enables efficient querying of formatted text data using commonly used query statements.
Owner:XIAN GRAPE CITY SOFTWARE CO LTD

A method for extracting fine-grained knowledge of a ship structure strength calculation formula

The application discloses a ship structure strength calculation formula fine-grained knowledge extraction method, which faces a ship structure strength calculation scene and comprises the following steps: generating a specified format text file by analyzing a ship structure strength calculation guide PDF; constructing a ship structure strength calculation formula fine-grained extraction large model; obtaining text-level full-quantity triple by adopting a general ontology model guided text-level full-quantity triple extraction method and a semantic-based triple edge relationship supplementing and correcting method; perfecting the ontology model by adopting an ontology model generation and attribute expansion method based on rule filtering and hierarchical clustering; formulating attribute value text description by adopting a triple snowflake type expansion method based on attribute semantic depth mining; and specifying attribute value reference description by adopting a triple supplementing method based on reference association expansion, so as to obtain formula-level full-quantity triple, thereby constructing a calculation formula graph and realizing the structured management of the calculation formula.
Owner:CHINA SHIP SCIENTIFIC RESEARCH CENTER

Electronic system and decision support method for a civil aircraft operator, including the development of aeronautical indicator(s), civil aircraft and associated computer program

Electronic system and method for assisting a civil aircraft operator in decision-making with the development of aeronautical indicator(s), civil aircraft and associated computer program. The present invention relates to an electronic system (15) for assisting a civil aircraft operator in decision-making (10), comprising: - an electronic device (18) for displaying information; - an electronic device (20) for developing aeronautical indicators, comprising: + a module (30) for acquiring at least one aeronautical information message comprising a header and a useful part, the useful part including several data fields, including a free-format text field, called the free field;+ a module (32) for calculating several aeronautical indicators for each aeronautical information message via the application of an artificial intelligence algorithm to said message, the artificial intelligence algorithm receiving the free field as input and delivering the indicators as output, these being distinct from the data contained in said message; + a module (34) for displaying a rendering of the aeronautical indicators. Figure for the abbreviation: Figure 1;
Owner:THALES SA

Intelligent question and answer method and system for documents in water taking permission field

The invention belongs to the field of natural language processing and large language models, and particularly discloses an intelligent question and answer method and system for a document in the field of water taking permission, and the method comprises the steps: analyzing a water taking permission document through layout recognition, and obtaining a json format text; performing knowledge extraction on the text, namely identifying a paragraph main entity through a paragraph main entity dual-channel detection and entity identification iterative training method, complementing the content of an indication pronoun in the text through long-distance anaphora resolution based on an elastic sliding window, and guiding a large language model to extract a triple by using knowledge extraction Prompt, data cleaning is carried out on the data; constructing a vector database index based on the cleaned triples, and when a user question is received, retrieving candidate triples related to the user question by using the vector database index; and inputting the retrieved candidate triples into the large language model to generate answers to the user questions. According to the method, the knowledge can be efficiently and accurately extracted and retrieved from the water taking permission document.
Owner:CHINA UNIV OF GEOSCIENCES (WUHAN)

Machine-readable constructional engineering standard digital processing method and device

The embodiment of the invention discloses a machine-readable building engineering standard digital processing method and a machine-readable building engineering standard digital processing device. A specific embodiment of the method comprises the steps of generating text sequence information according to a first preset processing mode; generating layout information and bibliography and reference relation information according to preset layout annotation information, the layout level identification model and preset bibliography annotation information; text information is generated according to preset text labeling information and a paragraph extraction model; according to a pre-trained formula identification model, generating special format text information; determining the text sequence information, the layout information, the bibliography and reference relation information, the text information and the special format text information as construction standard control information; and generating control parameter information according to the construction standard control information, and controlling the construction robot to perform construction processing according to the control parameter information. According to the embodiment, the accuracy of obtaining the building engineering standard information is improved, the construction time consumption is shortened, and the waste of equipment resources is reduced.
Owner:CHINA INST OF BUILDING STANDARD DESIGN & RES

Method and system for parsing a presentation document

The present disclosure provides a presentation document parsing method and system, wherein the presentation document parsing method comprises: obtaining recognition results of all pages of a presentation document; wherein the recognition results comprise intermediate state JSON data and Markdown format text content of all pages; obtaining a structured information extraction outline of the presentation document according to a user intention and the recognition results; determining key pages of the presentation document corresponding to the structured information extraction outline; inputting original page images, text recognition results and chart structure codes of the obtained key pages into a visual large model to output a presentation document structured parsing result conforming to the user intention. The present disclosure generates a structured information extraction outline of the presentation document according to the user intention, so that the user does not need to know the structure and content of the presentation document in advance, and the parsing path of the presentation document can be automatically planned, greatly reducing the use threshold.
Owner:HUA DATA TECH (SHANGHAI) CO LTD

Word document generation method and system, electronic equipment and storage medium

The invention provides a Word document generation method and system, electronic equipment and a storage medium. The method comprises the following steps: displaying a Web page containing a text editing area; obtaining input content and format requirements in the text editing area; judging whether the input content contains a structure identifier or not; if not, calling a semantic recognition model to perform structured analysis on the input content; if yes, performing hierarchical structure analysis on the input content according to a regularization rule and / or calling a semantic recognition model to perform structured analysis on the input content; and adjusting the input content according to the analysis result and the format requirement, and generating a Word document. Word document generation does not depend on a template, a fixed template does not need to be designed in advance, a text of a fixed structure does not need to be provided according to the template, a user can obtain the standard Word document by submitting the text of the free format through the Web page, operation is convenient and fast, universality is high, and flexibility is high.
Owner:HUA DATA TECH (SHANGHAI) CO LTD

Display method, device and storage medium of text to be displayed

PendingCN122369047AAlgorithmSlicing
This application relates to a method, apparatus, and storage medium for displaying text to be displayed. The method includes: upon acquiring the text to be displayed, acquiring the global offset of the target text within the text to be displayed that requires formatting adjustment; determining the global character range of each slice in the slicing results of the text to be displayed; determining, based on the comparison result between the global character range and the global offset, a target slice containing the character to be adjusted, and determining the local offset of the character to be adjusted within its respective target slice; and during the streaming display of the target slice of the text to be displayed, adjusting the format of the character to be adjusted corresponding to the local offset within the target slice before displaying it. This application solves the technical problem that formatted text cannot be displayed when displaying long text output by LLM.
Owner:GREAT WALL MOTOR CO LTD

Computer-implemented system and method for generating datasets that represent traffic scenarios

The present invention relates to a computer-implemented system (100) and a computer-implemented method for generating datasets (3) representing traffic scenarios, which are to be used for training and / or testing automated driving (AD) functions. The system (100) comprises one or more processors and a memory that stores one or more programs configured for execution by the one or more processors. The programs comprise at least one first large-language model (LLM) and / or vision-language model (VLM), which are trained on at least formatted text descriptions of traffic scenarios and on natural language to perform a task specified by a prompt. Furthermore, the programs include instructions for: • Transforming an initial data set (1) into a formatted text description of a corresponding real-world traffic scenario, • Generating an initial prompt (11) as input for the initial LLM or VLM acting as a modification agent (10), wherein the initial prompt (11) includes the formatted text description of the real traffic scenario and a user specification of a desired modification of the real traffic scenario, • Applying the first prompt to the modification agent (10) and receiving a modified formatted text description of a modified traffic scenario as output (12), and • Transforming the modified formatted text description into a corresponding data set (3) that represents the modified traffic scenario.
Owner:ROBERT BOSCH GMBH

Computer Implemented System and Method for Generating Data Sets Representing Traffic Scenarios

A computer implemented system and method for generating data sets representing traffic scenarios to be used for training and / or testing automated driving functions are disclosed. The system includes one or more processors and a memory that stores one or more programs that are configured to be executed by the one or more processors. The programs include at least one first Large Language Model (LLM) and / or Vision Language Model (VLM) being trained at least on formatted text descriptions of traffic scenarios and on natural language to perform a task given a prompt. Further, the programs include instructions to (i) transform an initial data set into a formatted text description of a corresponding real traffic scenario, (ii) construct a first prompt as input for the first LLM or VLM acting as modifier agent, wherein the first prompt includes the formatted text description of the real traffic scenario and a user specification of a wanted modification of the real traffic scenario, (iii) apply the first prompt to the modifier agent and obtain a modified formatted text description of a modified traffic scenario as output, and (iv) transform the modified formatted text description into a corresponding data set representing the modified traffic scenario.
Owner:ROBERT BOSCH GMBH

Print data processing device and printing system

The present invention provides a printing data processing device and a printing system that enable a user to directly print a table document through a handheld printer. It is characterized in that, after the content of the table document is displayed on the printing data processing device and the user selects the cells to be printed, the text content and text format of each cell are obtained from the table document through the cell content acquisition unit, and formatted text corresponding to each cell is generated through the formatted text generation unit. Further, the bitmap generation unit can generate a corresponding bitmap based on the formatted text and process it into a scrolling print content through the scrolling print content processing unit for the user to print. Therefore, through this printing system, as long as the user selects the cells to be printed in the table document, the handheld printer can print the table document, avoiding the need for the user to repeatedly set the format in the table document through other means and enabling direct printing to be achieved.
Owner:RICOH CO LTD

Knowledge graph-based learning path planning method and application

The invention discloses a knowledge graph-based learning path planning method, which comprises the following steps of: acquiring learning behavior data of a learner, and preprocessing audio data in the learning behavior data into formatted text data; analyzing the preprocessed text data, and performing entity extraction and relationship extraction to obtain a head entity, a relationship and a tail entity; the head entity, the relation and the tail entity form a triple, and an English knowledge graph of the student is formed through the triple to represent the oral English learning condition of the student; performing interactive fusion on entities and relation embedding of the English knowledge graph, and performing training by using an FIMA-GNN model to respectively obtain updated entities and relation embedding; and embedding based on the updated entities and relationships to obtain a target knowledge graph, and reasoning according to the target knowledge graph to generate a personalized learning path. According to the method, multi-modal data acquisition and analysis of oral English learning behaviors of learners can be realized, and personalized learning content suggestions are generated.
Owner:HUAZHONG NORMAL UNIV

Big model based vertical field data construction method

The application provides a large model-based vertical field data construction method, and belongs to the technical fields of data processing and artificial intelligence, and comprises the following steps: converting a vertical field source document into an intermediate format text and dividing the intermediate format text into multiple text blocks; inputting the text blocks into a pre-trained generative language model, guiding the pre-trained generative language model to generate a plurality of candidate questions according to the contents of the text blocks according to a pre-designed prompt word, and performing preliminary screening and fine screening on each candidate question to obtain a question set; predefining a mode of a knowledge graph according to the field characteristics of the vertical field, processing all the text blocks based on an information extraction model, and constructing a field knowledge graph; and performing local context retrieval based on the text blocks of the source of each final question in the question set, performing global knowledge retrieval based on the field knowledge graph, and generating a final answer and a final thinking chain. The application is suitable for different vertical fields, and can effectively improve the quality of data and ensure the accuracy of question generation.
Owner:PEKING UNIV

Meteorological data format conversion method based on data translation

The invention discloses a meteorological data format conversion method based on data translation, and relates to the technical field of data processing driven by artificial intelligence, and the method comprises the following steps: carrying out zoning space cutting and time window alignment on multi-source meteorological data to generate standardized input; analyzing the standardized data into time tags and geographic tags to form structured semantic description; and constructing cue words according to the structured tag to drive the large language model to output a fixed format text. According to the method, standardization of model semantic input is achieved through semantic structure modeling; and on the basis of a label-driven cue word construction and term specification mechanism, the professional output capability of a large language model in meteorological services is enhanced. The risk of model illusion and semantic deviation is effectively reduced, and the accuracy and service adaptability of automatic meteorological text generation are improved.
Owner:JIANGSU METEOROLOGICAL OBSERVATORY