Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

307 results about "Plain text" patented technology

In computing, plain text is a loose term for data (e.g. file contents) that represent only characters of readable material but not its graphical representation nor other objects (floating-point numbers, images, etc.). It may also include a limited number of characters that control simple arrangement of text, such as spaces, line breaks, or tabulation characters (although tab characters can "mean" many different things, so are hardly "plain"). Plain text is different from formatted text, where style information is included; from structured text, where structural parts of the document such as paragraphs, sections, and the like are identified); and from binary files in which some portions must be interpreted as binary objects (encoded integers, real numbers, images, etc.).

Voice generation method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of medical health, financial science and technology and the like, and discloses a voice generation method which comprises the following steps: constructing a multi-language voice synthesis model, obtaining plain text data and paired voice text data, and constructing an expansion vocabulary; updating a language perception embedding layer and model parameters, and converting an input text into a mark sequence; and the encoder extracts context semantic features, extracts pronunciation rule features, and the decoder fuses the features to generate an acoustic feature sequence, and converts the acoustic feature sequence into target voice data. According to the invention, the multi-language speech synthesis model is combined with the language perception embedding layer, so that the speech generation capability of a low-resource language is improved; the text conversion accuracy is improved by expanding the vocabulary, the target language learning ability is enhanced by unsupervised training, the low data environment adaptability is optimized by supervised training, and the speech naturalness and fluency are improved by feature fusion.
Owner:PING AN TECH (SHENZHEN) CO LTD

Retrieval augmented generative question and answer boosting

Systems or techniques are provided for facilitating retrieval augmented generative question and answer boosting. In various embodiments, a system can access a plain text question regarding a scientific instrument. In various aspects, the system can generate, via a large language model that references a document-graph repository, a structured or unstructured answer for the plain text question. In various instances, the document-graph repository can comprise a plurality of document-graphs that respectively correspond to a plurality of technical documents. In various cases, for a first document-graph that corresponds to a first technical document, leaf nodes of the first document-graph can represent respective text blocks written in the first technical document, and non-leaf nodes of the first document-graph can respectively represent a document title, one or more section headings, and one or more scientific instrument identifiers written in the first technical document and beneath which the respective text blocks are nested.
Owner:PPD DEVELOPMENT LP +2

Engineering document index consistency proofreading method and system based on multi-modal large model

The invention relates to an engineering document index consistency proofreading method and system based on a multi-modal large model, and the method comprises the steps: Q1. OCR detection and recognition: carrying out the optical character recognition and format analysis of a source document, converting an uploaded PDF document into a processable text message in a Markdown format, and carrying out the format discrimination of a table, a formula and a plain text; and Q2, table and formula processing: adopting a hierarchical processing strategy, intelligently selecting an optimal processing mode according to the complexity of the table, and converting table information into a descriptive long text through a language large model and cue words. According to the method, accurate, reliable and efficient document index checking service can be provided for a user, the quality and efficiency of professional document processing are remarkably improved, the efficiency and quality of knowledge graph construction are remarkably improved, a knowledge verification system capable of being evolved continuously is established, and the method is suitable for popularization and application. And a reliable technical support is provided for knowledge management and professional decision-making in a complex field.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

Multi-modal retrieval system based on similarity between image and text feature vectors

A multi-modal retrieval system based on the similarity between image and text feature vectors, which system belongs to the technical field of information retrieval. In the present invention, advanced deep learning technology is used to extract high-dimensional feature vectors of image data and text data, and the similarity between these feature vectors is computed by means of a specific similarity algorithm. The system can process various types of query inputs, which take the form of plain text, plain images and a combination thereof, thereby providing a flexible search mode.
Owner:INSPUR CLOUD INFORMATION TECH CO LTD

Data circulation trusted system and device based on block chain and privacy calculation

The invention discloses a data circulation trusted system and device based on a block chain and privacy calculation, and relates to the technical field of trusted data. Comprising a data supply end, a data demand end and a credible supervision end, the data supply end comprises a digital identity identification module which is used for generating a dynamic digital identity according to an enterprise unified social credit code and an ISO industry classification code and submitting the dynamic digital identity to a first block chain of the credible supervision end for authentication; the mixed desensitization module is used for carrying out first 3-bit plaintext reservation and hash replacement on the sensitive field by adopting an AES-256 algorithm, binding a UTC timestamp to generate a desensitization record and uploading the desensitization record to a first block chain; and the zero-knowledge proof generation module is used for converting the privacy calculation result into a zero-knowledge proof voucher and pushing verification parameters to a sandbox environment of the data demand end. According to the invention, the security and privacy of data in a circulation link and the credibility of transactions are guaranteed, and the problems of privacy leakage, uncredible transactions, difficult supervision and the like possibly existing in data circulation are solved.
Owner:GUIZHOU DIGITAL INNOVATION HLDG (GRP) CO LTD

Electronic acquisition of troubleshooting knowledge for medical imaging scanners

Systems or techniques that facilitate electronic acquisition of troubleshooting knowledge for medical imaging scanners are provided. In various embodiments, a system can access a plain text service manual and a service complaint that are associated with a medical imaging scanner. In various aspects, the system can identify, via named-entity recognition or natural language processing, a semantic hierarchy of the plain text service manual. In various instances, the system can record, via graphical user interface tracking, how a service technician that is troubleshooting the service complaint sequentially navigates through the semantic hierarchy, thereby yielding a troubleshooting trace that corresponds to the service complaint.
Owner:GE PRECISION HEALTHCARE LLC

Multi-dimensional storage system for cultural heritage archives

The invention discloses a cultural heritage file multi-dimensional storage system, which relates to the field of cultural heritage digital protection and intelligent storage and comprises a data acquisition and processing module, a double-library management module, a multi-dimensional index retrieval engine, a security and convergence module and a backup and disaster recovery module. According to the method, structured texts and unstructured files are uniformly accessed through various interfaces, a special tool chain is adopted for format standardization, audios and videos are transcoded into an H.264 / AAC format through FFmpeg, images are subjected to resolution adjustment and compression through ImageMagick, plain texts are extracted from texts through ApacheTika and are converted into UTF-8 codes, standardized data are uniformly converted into an intermediate format, and the format of the intermediate format is converted into the intermediate format. According to the technical scheme, the method comprises the following steps of: standardizing input data by adopting a three-layer duplicate removal strategy, facilitating other modules to receive data, performing refined redundancy processing on the standardized input data by adopting a three-layer duplicate removal strategy, remarkably reducing storage redundancy, partitioning a cultural heritage data main library according to three layers of ''project type-medium type-age'', and partitioning an inheritor information library according to two levels of ''province-inheritor number prefix''.
Owner:济宁市退役军人服务中心

Complex text OCR (Optical Character Recognition) error recognition and repair method based on large language model

The invention discloses a complex text OCR (Optical Character Recognition) error recognition and repair method based on a large language model, and relates to the technical field of text processing, the method comprises the following steps: step 1, a text preprocessing module recognizes and excludes a non-text area in a preliminary text result generated by OCR to ensure that residual content is pure text input, and a text preprocessing module is used for preprocessing the residual content; obtaining a text result of the plain text data; step 2, an error recognition module performs entity recognition on a text result of the plain text data by using a large language model, and provides revision suggestions for recognized suspected errors; 3, the database provides knowledge recall for the recognized entities and errors, and supplementation is provided for follow-up repair; and 4, correcting errors in a text result of the plain text data by an error repairing module through the acquired context information and supplementary data provided by the database, and outputting a document.
Owner:数字宁波科技有限公司

Updatable attribute condition proxy re-encryption method with forward security and quantum attack resistance

The invention provides an updatable attribute condition proxy re-encryption method with forward security and capable of resisting quantum attacks. The method comprises the steps that an authorization manager generates and publishes public parameters; the entrusting party and the entrusted party generate a public and private key pair of the entrusting party and the entrusted party according to the public parameters; the entrusting party encrypts the plaintext to generate an encrypted ciphertext and sends the encrypted ciphertext to the cloud server; the entrusting party generates an updated public key for the entrusting party according to the public key of the entrusting party, and the updated public key is used for updating the ciphertext of the private key of the entrusting party; the entrusting party generates a re-encryption key associated with the control strategy; the cloud server re-encrypts the encrypted ciphertext, generates a re-encrypted ciphertext and sends the re-encrypted ciphertext to the trusted party; the trusted party generates an updated private key; and the trusted party decrypts the re-encrypted ciphertext by using the updated private key. According to the updatable attribute condition proxy re-encryption method, public and private keys of a receiver are allowed to be periodically alternated based on a key asynchronous updating mechanism, so that forward security is realized; and fine-grained control of ciphertext conversion is realized through an attribute control structure.
Owner:JINAN UNIVERSITY

Memory graph query engine with persisted storage

Various examples of improving an in-memory graph query engine using a persisted storage component are provided. The method includes updating data stored in an in-memory graph query engine and, based on updating the data, converting the data to a plain text form that may be more efficiently stored in the persistent storage component. The method further includes updates to additional in-memory graph query engines from the persistent storage component such that in-memory data stored in the graph query engines is synchronized.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Systems and methods for providing explainability of natural language processing

Systems or techniques that facilitate systems and methods for providing explainability of natural language processing are provided. In various embodiments, a system can access a plain text clinical sentence. In various aspects, the system can generate, via execution of a first machine learning model, an assertion status classification label for a word of interest in the plain text clinical sentence. In various instances, the system can extract, from a hidden attention layer of the first machine learning model, word-wise attention scores corresponding to the plain text clinical sentence and render, on an electronic display, both the assertion status classification label and a graphical representation of the word-wise attention scores. In various cases, the system can determine, via execution of a second machine learning model, a reliability score for the assertion status classification label, based on the word-wise attention scores, and can render the reliability score on the electronic display.
Owner:GE PRECISION HEALTHCARE LLC

Encryption Key Distribution System

Customers of a software platform, such as a unified communications as a service platform, are enabled to control their own encryption keys used to encrypt and decrypt data from various communication services in the software platform. A key connector service is used to coordinate communications with a key management server for generating plaintext keys for data encryption and encrypted keys based on the plain text keys, and with a key broker server for storing and retrieving copies of the encrypted keys. A context identifier may be associated with an encrypted key and used to track which data has been encrypted with an underlying plaintext key. Examples of data encrypted may include conference recordings, webinar recordings, phone call recordings, voicemails, emails, and calendar tokens.
Owner:ZOOM COMMUNICATIONS INC

Large visual language model illusion mitigation method and device

The invention belongs to the technical field of artificial intelligence and multi-modal large models, and particularly relates to a large visual language model illusion relieving method and device. The method comprises the following steps of: acquiring a complete visual token of an original image and a text token of a text prompt, connecting the tokens and jointly inputting the tokens into a large language model decoder; calculating an attention score matrix of the text token and all the visual tokens based on a cross-modal dynamic sampling strategy so as to sample key visual tokens; obtaining classification tokens of the original image, and screening significant visual tokens based on the classification tokens and attention scores of the visual tokens in the complete visual tokens; and carrying out adaptive attention enhancement on the significant visual token and the key visual token, and subtracting the logs distribution influence of plain text input from the logs distribution with enhanced visual information by comparing decoding strategies so as to obtain final target text output. The objective of the invention is to alleviate illusion problems in large-scale visual language models.
Owner:HANGZHOU PENGUIN TECH CO LTD

Intelligent extraction method and system for input text containing mathematical formula

The present invention belongs to the technical field of text processing, and relates to an intelligent extraction method and system for an input text containing a mathematical formula. The method comprises: 1) performing format determination, conversion and preprocessing on an input text; 2) performing angle correction on the preprocessed image-formatted text; 3) performing formula detection; 4) performing layout analysis; 5) for inline formulas, according to a formula detection box, determining whether a corrected OCR detection box contains an inline formula, and splitting the OCR detection box containing the inline formula, so as to obtain a text-only OCR detection box; 6) performing formula recognition to obtain a formula recognition result; 7) performing text recognition to obtain a text recognition result; and 8) on the basis of a layout analysis box and a layout category thereof, performing same-line detection box determination and merging on the formula recognition result and the text recognition result, so as to obtain an extraction result for the input text. The present invention can effectively improve the extraction efficiency and accuracy for input texts containing mathematical formulas.
Owner:BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD

Connectionless-virtual private network for secure cloud to user communication over the internet using a plurality of servers

The disclosure provides a system / method / scheme to securely send data from a cloud, or cloud service provider, to users via a secure connectionless system, referred to herein as a C-VPN communication infrastructure (C-VPN CI). In one example a method of communicating from a cloud service provider to a user via a C-VPN CI includes: (1) obtaining, by a cloud service provider, security parameters from a SDE Cloud server operating on a computing system of the cloud service provider, wherein the security parameters include a set of mathematical rules and values for converting plain text to ciphertext, (2) creating a secure communication using the security parameters received from the SDE Cloud server, wherein the secure communication includes a secure header and secure data, and (3) sending the secure communication to the user via a generic electronic message delivery system.
Owner:TALATI FAMILY

Contract risk auditing method and device

The invention relates to the field of artificial intelligence, and particularly provides a contract risk auditing method and device, and the method comprises the following steps: S1, splitting a contract content paragraph, and decomposing a contract text into a plurality of independent paragraph units with a logic relation; s2, converting the contract text into a structured data unit, and performing text analysis and risk factor extraction; s3, table textualization: converting table data in the contract into a plain text format which can be processed by a large model; s4, constructing a list table to systematically organize key information in the contract; s5, performing text semantic analysis to deeply understand the contract text; s6, element checking: carefully reviewing and verifying key elements in the contract text; and S7, risk information summarization: summarizing and analyzing potential risks in the contract text. Compared with the prior art, the method has the advantages that potential risks in the text can be recognized and predicted through learning of a large amount of text data, and therefore the efficiency and accuracy of contract risk auditing are remarkably improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

An electronic binder system (ebinder) for processing source data to EDC systems

The present invention provides a method and system for automatically and seamlessly processing clinical trial source data into electronic data capture (EDC) systems. In one embodiment, a file structure is defined for an electronic binder system (eBinder); source data is uploaded to the eBinder; the source data is encrypted, Patient Identifiable Information in the source data is masked; the source data is converted into machine readable plain text in the JavaScript Object Notation (JSON) format using Natural Language Processing (NPL) technologies; the JSON data is converted into tabulated machine readable data in the HyperText Markup Language (HTML) format using NPL technologies; the HTML data is converted into machine understandable data using NPL technologies; the machine understandable data is populated into EDC datasets using NPL technologies; the source data and converted data are displayed side-by-side for source data verification; and a platform is provided for regulatory data verification or auditing.
Owner:XIE TAI +1

Open domain-oriented adaptive public opinion data classification method and system

The invention discloses a self-adaptive public opinion data classification method and system for an open domain, and the method comprises the steps: collecting original text data from a plurality of data sources, carrying out the preprocessing of the original text data, obtaining a plain text list, converting the plain text list into high-dimensional semantic vectors in batches, and enabling all high-dimensional semantic vectors to form an embedded matrix; calculating a minimum clustering number and a maximum clustering number according to the number of texts in the plain text list, generating various clustering schemes corresponding to the embedded matrix through all clustering thresholds in a clustering range, calculating a comprehensive score of each clustering scheme, and selecting the clustering scheme with the highest comprehensive score as an optimal scheme; generating subject terms of all clustering clusters based on a large model; and integrating the optimal parameters, the texts in each cluster and the subject terms in each cluster, and outputting a structured classification result. The open domain public opinion data can be efficiently, intelligently and interpretably classified.
Owner:CHENGDU SPACEON IND CO LTD

Title level identification large model training method, title identification method, system and program product

According to the title level identification large model training method, the title identification method and system and the program product provided by the invention, the first title information covering all texts of the original file is constructed through the multi-modal semantic model, and the limitation of plain text identification is made up in combination with the element matching page picture; second title information with semantic and visual features is generated through multi-modal fusion, so that the title judgment accuracy is improved; in the training process, effective title objects are screened in combination with original title objects to optimize pre-training data, and a large model which is high in precision and adapts to complex scenes is cultivated; in the identification process, the title information to be identified is constructed based on the valid title object. And performing dynamic branch processing according to a calling condition, if not, directly outputting an answer, if yes, generating accurate final title information by means of a trained large model, and finally constructing and outputting the answer, thereby realizing training and identification full-link coordination, considering complex scene adaptability and efficient and accurate identification, and comprehensively improving the structuralization and practicability of document title identification.
Owner:SHANGHAI HUNDSUN JUYUAN DATA SERVICE CO LTD +1

Rich text data transmission method, electronic device and program product

The invention provides a rich text data transmission method, electronic equipment and a program product. The rich text data transmission method comprises the steps of obtaining a plain text and rendering information of the plain text included in to-be-transmitted rich text data; obtaining rendering description data of the rendering information based on the plain text; generating structured rendering information based on the rendering description data; and transmitting the plain text and the structured rendering information to a receiving end to execute transmission of the rich text data to be transmitted.
Owner:KE COM (BEIJING) TECHNOLOGY CO LTD

System for synchronously displaying map in combination with content information

The invention provides a system for synchronously displaying a map in combination with content information. The system comprises a content information module, a geographic information detection module, a map reading module, a display control module and a display which are connected in sequence, the content information module is used for providing content information to be displayed; the content information comprises at least one of plain text, image-text, video and audio; the geographic information detection module is used for determining geographic information to be displayed from the content information; the map reading module is used for acquiring a map corresponding to the geographic information to be displayed; and the display control module is used for controlling the display to display the content information and the map. According to the method and the device, the content information and the related map can be synchronously displayed, and the user can realize the understanding of the map information, improve the cognition of the world and improve the user experience on the basis of wide contents and wide places after using the content information and the related map for a large amount of time.
Owner:SHANGHAI ROSE-EARTH CULTURAL TECHNOLOGY CORP LTD

Hybrid retrieval method, system and equipment based on multi-algorithm fusion and storage medium

The invention provides a mixed retrieval method, system and device based on multi-algorithm fusion and a storage medium, and the method comprises the steps: obtaining various types of document data, preprocessing the document data, and obtaining plain text data; segmenting the plain text data into a plurality of pieces of block node data, and generating a plurality of derivative problems for each piece of block node data by using a large language model; carrying out vectorization coding on the block node data and the derivation problem, and storing the block node data and the derivation problem in a vector database; receiving a user query question, and generating a plurality of related preset questions for the query question by using the large language model; converting the query question and the preset question into a vector form, executing keyword retrieval and vector retrieval, and obtaining corresponding candidate results in a vector database; and carrying out weighted fusion on the candidate results through a weighted reciprocal rearrangement algorithm to obtain most relevant retrieval data. According to the method, the efficiency of information retrieval and the correlation of results can be improved, and the retrieval requirement of a user in a complex semantic scene is met.
Owner:XIAMEN INTRETECH

SysML model generation method based on large language model in conceptual design

The invention provides a SysML model generation method based on a large language model in conceptual design, which comprises the following steps of: guiding LLM by utilizing a thinking chain cue word template, and converting user input into a conceptual design scheme described in a plain text form; utilizing a cue word template I to guide LLM to split the conceptual design scheme into a sentence list containing complete semantics; using a cue word template 2 to guide LLM to extract entities and relationships thereof from each sentence, and outputting a triple; based on the triple, creating a SysML model hierarchical structure according to a predefined rule; and creating an XMI file by following an XMI format specification to obtain a complete SysML model. According to the method, the automation degree and generalization of SysML model generation are remarkably improved, less text input can be effectively processed, a conceptual design scheme model which can be directly used for actual system design and adapts to products is generated, and the method is particularly suitable for high-complexity large-scale system design.
Owner:ZHEJIANG UNIV

Multi-source document management method and device based on knowledge construction and fusion storage

The invention discloses a multi-source document management method and device based on knowledge construction and fusion storage, and relates to the technical field of document management, and the method comprises the steps: receiving a multi-source document to an object storage system, and recognizing the document type; performing document analysis and structure extraction according to the document type; extracting pictures in the image-text mixed content, uploading the pictures to an object storage system, generating a mapping dictionary, and inserting picture marks in a text part of the mapping dictionary; carrying out structured processing on contents of the table key value pairs to extract summaries and abstracts; performing standardization processing and semantic slicing on the processed image-text mixed content, table key value pair content and / or plain text content to generate knowledge fragments; constructing a knowledge extracting questions from knowledge fragments; the knowledge fragments are stored in a first index, and the questions and IDs of the corresponding knowledge fragments are stored in a second index; generating vectors for each knowledge fragment and question, embedding and writing the vectors into a vector database; and performing document management based on the knowledge base. The document management efficiency is improved.
Owner:DIGITAL CHINA SYST INTEGRATION SERVICE

Industrial automation design environment prompt engineering for generative AI

An integrated development environment (IDE) for designing, programming, and configuring aspects of an industrial automation system uses a generative artificial intelligence (AI) model and associated neural networks to generate portions of an industrial automation project in accordance with functional requirements provided to the industrial IDE system in intuitive formats, such as spoken or written plain language text. The system uses generative AI to translate plain language requests or functional specifications into industrial control code, human-machine interface (HMI) applications, device configuration settings, or other aspects of an industrial control project.
Owner:ROCKWELL AUTOMATION TECH INC

PDF text extraction method and system capable of maintaining text reading sequence, and device

A PDF text extraction method and system capable of maintaining a text reading sequence, and a device. The method comprises: for a PDF text type, performing text extraction to obtain text block coordinate information; on the basis of the text block coordinate information, performing coordinate alignment and calculating coordinates, a row number row_no, a column number col_no, a row processing sequential number row_sequential_number and a column processing sequential number col_sequential_number of each text block; and on the basis of a text recovery type and calculation results, performing text recovery to obtain plain text content.
Owner:FUJIAN FOXIT SOFTWARE DEV LTD

A financial statement automatic analysis system based on large language models

The present invention discloses a financial statement automatic analysis system based on a large language model, comprising: a financial analysis DSL generation large model, a financial analysis DSL filter, and a financial analysis DSL decoder; when the system receives at least one financial statement analysis request in the form of natural language, the financial analysis DSL generation large model automatically generates a primary financial analysis DSL script, inputs the primary script into the financial analysis DSL filter to obtain a complete financial analysis DSL script, and then inputs the complete script into the financial analysis DSL decoder for decoding to obtain a prompt text containing financial analysis results and knowledge text, and the inference model returns the financial analysis and explanatory notes in plain text form to the user interface. The financial statement automatic analysis system based on the large language model of the present invention realizes the automatic generation of financial analysis reports for natural language requests, significantly improving the efficiency and accuracy of financial statement analysis.
Owner:ZHEJIANG UNIV

Automatic standard file classification method based on artificial intelligence

The invention discloses a standard file automatic classification method based on artificial intelligence, and relates to the technical field of text classification, and the method comprises the steps: obtaining original data of a to-be-classified file, and carrying out the analysis and preprocessing of the original data of the to-be-classified file, and obtaining metadata, chapter structure information and plain text content; inputting the plain text content and the chapter structure information into a multi-granularity semantic pyramid model for analysis and fusion, and generating a document feature vector; constructing a standard classification system knowledge graph by using the standard classification system data to obtain a classification name mapping table, and performing reinforcement learning on the standard classification system knowledge graph by using a graph neural network technology to generate an enhanced feature vector; and calculating the semantic similarity between the document feature vector and the enhanced feature vector. According to the method, hierarchical reasoning is performed in the knowledge graph based on the similarity to obtain the target classification code, and the classification result is matched and output through the mapping table.
Owner:CHINA STANDARD TECH DEV CORP

Description generation method based on continuous zero sample

The invention provides a description generation method based on continuous zero samples. The description generation method comprises the following steps: sequencing plain text training corpora according to the difficulty of continuous learning; setting that each training step t only can access the t-th corpus, and for the t-th learning step, generating a synthetic image for an input text of the t-th corpus by using a diffusion model; extracting image embedding of the synthesized image by using a CLIP image encoder; retrieving a group of text description embedding with similar semantics in a text corpus by using image embedding; image embedding and text description embedding are input into a fusion module network, and the output of a fusion module is defined as a soft prompt of a large language model; for an input text, key entities are extracted from the text through word frequency statistics so as to construct a soft prompt and a hard prompt of the large language model; hard prompts and soft prompts are input into a large language model network, and key parts in input are concerned through an attention mechanism.
Owner:JIANGSU VOCATIONAL & TECHNICAL UNIVERSITY OF ARCHITECTURE

Risk early warning question answering method and device based on large language model

The invention provides a risk early warning question answering method and device based on a large language model, and relates to the technical field of data processing, the method comprises the following steps: converting different types of specialized documents into plain text data, and carrying out text partitioning processing to obtain a plurality of knowledge data blocks; vectorizing a risk early warning problem input by a user, performing similarity matching with a stored vector database, retrieving knowledge data blocks of which the similarity meets a preset condition from the plurality of knowledge data blocks, and constructing optimized prompt data in combination with the risk early warning problem; inputting prompt data to a preset first model, and obtaining reply output of the preset first model; and training a preset second model by adopting the prompt data and the reply output in combination with an existing question and answer data set, and obtaining a question and answer model for risk early warning after training is completed. According to the method and the device, the user risk early warning problem can be accurately replied in combination with industry professional knowledge and variable data.
Owner:WUHAN SHENGHUAWEIYE TECHNOLGY CO LTD