Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

67 results about "Document recognition" patented technology

Intelligent document recognition is a new technology that promises to transform the way businesses handle document processing. An Intelligent document recognition system analyzes the content of the document that it receives, and looks for certain keywords that match its database of business terms.

Document data entry method, electronic device, storage medium and program product

The invention discloses a document data entry method, electronic equipment, a storage medium and a program product, and relates to the technical field of document recognition, and the document data entry method comprises the following steps: obtaining a to-be-recognized document containing at least one to-be-recognized page; the document to be recognized is recognized through the optical character recognition technology, a first recognition result and a target page are obtained, and the target page is a page to be recognized containing non-text elements; identifying the target page through the target large language model to obtain a second identification result; and inputting the first recognition result and the second recognition result into the target form according to the similarity between the form field of the preset target form and the first recognition result and the second recognition result, and obtaining the input target form. Through cooperative work of the OCR and the large language model, synchronous and efficient recognition of text and non-text information is achieved, and the accuracy and the automation level of document recognition and document data entry are improved in combination with an intelligent matching entry mechanism.
Owner:YILINYUN (SHENZHEN) TECH CO LTD

Complex document recognition method based on layout analysis and OCR (optical character recognition) and medium

The invention discloses a complex document recognition method based on layout analysis and OCR (optical character recognition) and a medium. Obtaining a to-be-recognized document image in real time, and performing preprocessing operation through a preprocessing module to obtain a to-be-recognized standard document image; processing the to-be-identified standard document image through a layout analysis module, and obtaining at least one target area image, a target area image type and an area vision matrix feature in combination with a layout analysis strategy method; performing text analysis on each target area image through a multi-mode OCR engine module according to the type of each target area image to obtain each text semantic analysis result and an area semantic matrix feature; obtaining a current visual text fusion feature through a multi-modal attention fusion module; and generating and feeding back a complex document identification output result according to the current visual text fusion feature. The problem that the accuracy is poor due to the fact that irregular layouts cannot be processed is solved, and the accuracy and flexibility of complex layout recognition are improved.
Owner:ZHEJIANG BAORONG MEDIA TECH (ZHEJIANG) CO LTD

Software development-oriented security processing method and device, equipment and medium

The invention relates to the technical field of data security, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a security processing method, device and equipment oriented to software development and a medium. Obtaining an architecture design document to identify potential safety hazards and generate design improvement suggestions; generating a code based on the business logic description and the design improvement suggestion, and completing security detection and repair to obtain a processed code and a code repair record; performing security test on the processed code and recording a test result; collecting data of exception identification, hidden danger identification, code detection and repair and security test to update the security knowledge base; and generating a security analysis report based on the demand exception list, the design improvement suggestion, the code repair record and the test result. According to the invention, through a security identification and restoration process from demand to test, early discovery of security problems, linkage processing and knowledge self-updating are realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Visual language large model-based document identification and structuring method

The invention discloses a document recognition and structuring method based on a visual language large model. The method comprises the following steps that S1, an input access layer receives a PDF or image document, and a page preprocessing and rendering layer unifies the resolution ratio and geometric parameters and performs denoising; s2, the heterogeneous recognition layer calls multiple recognizers in parallel on the same page to generate candidate results; s3, the alignment and fusion layer completes spatial alignment and text consistency evaluation of the candidate segments in a unified coordinate system to form a single main result; according to the method, the heterogeneous recognizers are called in parallel, a comprehensive scoring mechanism of space alignment, text consistency and model reliability is combined, fusion judgment is carried out on multi-source candidate results, segmentation errors and recognition deviation of a single model in complex scenes of nesting tables, cross-column titles and scanning noise can be avoided, and the recognition accuracy of the multi-source candidate results is improved. Stable output can still be kept in a multi-template and high-noise environment, and semantic consistency of recognition results is improved.
Owner:SHENZHEN SHENGWEI THREAD TECHNOLOGY CO LTD

Method and device for improving RAG recall effect

The invention provides a method and device for improving an RAG recall effect, and belongs to the technical field of computers, and the method comprises the following steps: file input and typesetting structure analysis: identifying a typesetting unit for an input file, extracting text content in the typesetting unit, and generating associated data of typesetting and content; constructing a two-dimensional relation graph: performing semantic segmentation based on the typesetting units, and extracting a logic relation of the typesetting units; constructing a two-dimensional relation graph of the semantic relation and the typesetting relation; and multi-dimensional information fusion recall: receiving user query and performing semantic analysis, recalling similar semantic slices from a semantic community and a typesetting community, executing double-graph cross validation, dynamically adjusting weights, and calculating and obtaining a final recall result. According to the method, the knowledge base construction mode of the RAG is optimized from the perspective of typesetting, the multi-dimensional relation between text semantics and typesetting logic is fused, information association in the knowledge base is more comprehensive, and the retrieval recall can be based on the semantic similarity and the typesetting logic at the same time, so that the recall effect is remarkably improved.
Owner:KYLIN CORP

Ocr-based medical document intelligent recognition method and system

The application discloses an OCR-based medical document intelligent identification method and system, relates to the technical field of document identification, and comprises the following steps: collecting a medical document image through a mobile terminal, generating a binary image based on the collected image through a U-Net combined with multi-scale feature fusion and an attention mechanism, and performing clipping; based on the clipped binary image, extracting text information through an OCR model, and extracting structured fields according to the layout rules and geometric distribution of the medical document; and performing deterministic rule judgment and risk assessment on the structured data. The application reduces the image calculation complexity through standard gray scale processing, improves the text region feature extraction accuracy through the U-Net structure combined with the CBAM attention mechanism, realizes effective fusion of multi-scale features and noise suppression in combination with the Attention Gate, realizes text direction correction in combination with the Hough transform, and improves the text detection robustness and recognition accuracy.
Owner:SHALLBRIGHT HEALTHTECH CO LTD

PDF (Portable Document Format) document content identification method and device, equipment and storage medium

The invention provides a PDF (Portable Document Format) document content identification method and device, equipment and a storage medium. The method comprises the following steps: acquiring an access link of an unanalyzed document; downloading the target PDF document from the object storage service according to the access link; when the content region corresponding to the page type in the target PDF document is a text region, dividing the content region into a text region, a table region and an image region according to the page type corresponding to each content region; and when the content region is a text region, extracting a native text character sequence from the text region, and performing similarity calculation on the native text character sequence to generate a semantic coherent paragraph text. According to the method, the sentences with similar semantics in the text region are automatically divided into the same text block based on the cosine similarity, so that the text content with coherent and complete semantics is analyzed from the document, the problem that text paragraphs are broken after document recognition is solved, and the semantic coherence of the document content is effectively improved.
Owner:SHENZHEN ISSMART SCI & TECH CO LTD

An evidence chain-based verifiable large model retrieval enhancement generation system and method

ActiveCN121301554BWeb data indexingBiological modelsShardFuzzy query
The application relates to the technical field of natural language processing, in particular to an evidence chain-based verifiable large model retrieval enhancement generation system and method, which comprises the following steps: receiving an initial query, identifying the ambiguity and information gap of the initial query in combination with associated retrieval documents, and generating a supplementary query set; generating a cited candidate answer based on the initial query, the supplementary query and the corresponding retrieval document, and verifying the information supportability of the cited candidate answer; extracting supportable information from the verified candidate answer and constructing a hierarchical attribution mapping relationship; integrating the above information to form to-be-evaluated information; if the to-be-evaluated information currently verified satisfies a preset sufficiency condition, integrating a generated preliminary answer and the to-be-evaluated information to synthesize a target answer. The application effectively solves the defects of the traditional RAG system, such as information integration fragmentation, one-sided query answering, inefficient attribution and excessive dependence on retrieval content, and has the advantages of comprehensive answer, verifiability and lightweight deployment.
Owner:JIANGNAN UNIV +2

Document identification method and device

The invention discloses a document identification method and device, and relates to the field of artificial intelligence. The method comprises the following steps: acquiring a to-be-identified document; performing optical character recognition on the document to obtain a first recognition result; performing multi-mode identification on the document to obtain a second identification result; and performing text synthesis on the first recognition result and the second recognition result to output a structured recognition result. According to the method, the OCR and the multi-modal recognition are combined, compared with single OCR, the recognition accuracy is higher, and the structured recognition result can be output through text synthesis.
Owner:CHINALCO DIGITAL (CHENGDU) TECHNOLOGY CO LTD

Foreign affairs certificate recognition method and system based on multi-modal fusion

This application relates to the field of image recognition technology, specifically to a method and system for recognizing diplomatic documents based on multimodal fusion. The method includes: using semantic segmentation to obtain text recognition regions and portrait regions from diplomatic document images; performing edge detection and rectangle fitting on the text recognition regions to determine the text ratio of each fitted rectangle and the text adhesion degree of each image block; matching pixels on both sides of the portrait region, analyzing the symmetry of each matching point with respect to the facial symmetry line of the portrait region, and determining a local enhancement factor; determining the image gain requirement based on the pixel ratio of the text recognition region to the portrait region in each image block, and correcting the contrast cropping threshold of the image enhancement algorithm to enhance the diplomatic document image; and recognizing the text content and portrait content in the enhanced diplomatic document image. This application aims to enhance diplomatic document images and increase the reliability of diplomatic document recognition.
Owner:BEIJING DEXUN AVIATION SERVICE CO LTD

Unstructured document recognition method and system

This invention relates to the field of document recognition technology, and more particularly to a method and system for recognizing unstructured documents. The method includes: acquiring a PDF document to be recognized; segmenting the PDF document to form a row dataset; dividing the row dataset into regions to obtain text regions and table regions; and using a recognition method to recognize data in the text regions and table regions respectively, obtaining data in the text regions and data in the table regions respectively. The purpose of this invention is to solve the problem of low accuracy in recognizing text and tables in unstructured documents using existing technologies.
Owner:CRRC QINGDAO SIFANG CO LTD

A PDF document content recognition method, device, equipment and storage medium

The application provides a PDF document content recognition method and device, equipment and a storage medium, the method comprises the following steps: obtaining an access link of an unanalyzed document; downloading a target PDF document from an object storage service according to the access link; when the content area corresponding to the page type in the target PDF document is a text area, the content area is divided into a text area, a table area and an image area according to the page type corresponding to each content area; when the content area is a text area, the original text character sequence is extracted from the text area, and the semantic coherent paragraph text is generated by similarity calculation on the original text character sequence. Based on the cosine similarity, the application automatically divides the sentences with similar semantics in the text area into the same text block, so that the document analysis obtains semantic coherent and complete text content, solves the problem of broken text paragraphs after document recognition, and effectively improves the semantic coherence of the document content.
Owner:SHENZHEN ISSMART SCI & TECH CO LTD

Document identification method and computer program product

PendingCN121808227ASolve the problem of insufficient recognition accuracyeasy to handleNatural language data processingDocument recognitionRapid processing
The invention discloses a document recognition method and a computer program product, and relates to the technical field of data processing.The document recognition method comprises the steps that the document complexity of a target document is recognized, and a text analysis strategy matched with the document complexity is determined; the text analysis strategy comprises at least one of a lightweight model analysis strategy, a large model analysis strategy and a lightweight collaborative large model analysis strategy; and analyzing the target document according to the text analysis strategy to obtain a document identification result of the target document. According to the method, the problem that the recognition precision of the mixed elements in the complex document is insufficient is effectively solved, low-complexity content can be rapidly processed locally by establishing a complexity evaluation and strategy matching mechanism, high-complexity content can obtain sufficient computing resources, efficiency and precision of medium-complexity content are balanced through a collaborative mechanism, and the recognition precision of the mixed elements in the complex document is improved. Therefore, the accuracy and the real-time performance of recognizing the complex document are greatly improved, and meanwhile, the resource utilization rate of the edge end and the cloud end is also greatly improved.
Owner:HANGZHOU ROBAM APPLIANCES CO LTD

Customs declaration document identification method and device

This invention provides a method and apparatus for recognizing customs declaration documents, relating to the field of artificial intelligence technology. The method includes: acquiring a target image of a customs declaration document to be recognized; inputting the target image and prompt words into a visual language model to obtain a first recognition result output by the visual language model; the prompt words are generated based on preset business requirements; the visual language model is obtained by performing low-rank adaptive fine-tuning of a CogVLM40B model; performing optical character recognition on the target image to obtain a second recognition result; and correcting the first recognition result based on the second recognition result to obtain a target recognition result corresponding to the customs declaration document. The customs declaration document recognition method provided by this invention improves the efficiency and accuracy of customs declaration document recognition by integrating the semantic understanding advantages of visual language models and the character accuracy advantages of optical character recognition through automated collaborative verification.
Owner:SINOTRANS +1

Method and system for metadata extraction for document identification

A method includes obtaining a document comprising earth property data regarding a geological region of interest, preprocessing the document to form at least one preprocessed document and determining, using a set of trained machine-learned models processing the at least one preprocessed document, a category of the document and a type of the document. The method further includes determining, using a natural language processing algorithm, metadata attributes of the document and a title of the document, and updating a database storing the document with the title, the category, the type and the metadata attributes. The method further includes identifying, by a planning module processing a query, the document from the database based on at least one of the title, the category, the type and the metadata attributes and planning a wellbore path in the geological region of interest using the earth property data comprised in the document.
Owner:SAUDI ARABIAN OIL CO

Artificial intelligence-based document recognition management system

The AI-based bill identification management system of the present application belongs to the field of artificial intelligence technology, and comprises a candidate paradigm generation unit for collecting image data of bills to be processed and extracting multi-modal features from the image data; the multi-modal features include text content, position coordinates and layout features; the candidate paradigm generation unit is also used to perform semantic analysis on the multi-modal features by using a pre-trained large language model to generate candidate paradigms and calculate the paradigm reasoning confidence of the candidate paradigms; a consistency checking unit is used to generate instantiated data according to the candidate paradigms and compare the instantiated data with a pre-constructed historical business knowledge graph to calculate a cross-document checking rate; the present application can stably process bills of various formats, achieving a substantial breakthrough in technology.
Owner:NINGBO JUXUAN INFORMATION SCI & TECH

Intelligent supervision cabinet based on cooperative movement and method thereof

The invention discloses an intelligent supervision cabinet based on collaborative movement and a method thereof, and belongs to the technical field of intelligent storage. An operation hole is formed in the front end face of the box; self-closing type door assemblies with the functions of closing locking, complete opening locking, automatic closing and taking and placing position indicating are arranged in the operation hole from top to bottom at equal intervals, and the adjacent ends of the adjacent self-closing type door assemblies are in attached sliding connection. A rotary storage assembly is connected in the box body, and the self-closing door assembly is connected with the rotary storage assembly; a movable image acquisition assembly is also connected in the box body; the front side wall of the box body is connected with a voice broadcast assembly, a camera shooting assembly, a touch screen assembly, a scanning assembly used for scanning a two-dimensional code on a deposit bill or a material taking bill, a fingerprint recognition assembly and a certificate recognition assembly. According to the invention, a many-to-one effect is realized, and a complete automatic checking closed-loop system specially designed for the closed intelligent cabinet is formed.
Owner:SHANGHAI BINGYAO TECHNOLOGY CO LTD

Cloud resource control method based on big data

The invention discloses a cloud resource control method based on big data, which comprises the following steps: S1, a user downloads software stored with the cloud resource control method, takes effect after authorization of the user, and extracts and monitors information in a mobile phone of the user; s2, in a chat interface of the instant messaging tool, after a user logs in the instant messaging tool by using account information on the mobile terminal, triggering the entered chat interface; s3, monitoring related information of a user chat interface in a monitoring time period, and when the group chat is positioned to belong to the group message hot area, starting a record analysis module; s4, when the coverage rate sent by the document obtained by the system exceeds the document recognition standard rate, analyzing the browsing situation of a member who clicks the document to browse the document, and judging whether the document needs to be stored or not, the monitoring accuracy of a control system and the control density of cloud resource utilization are improved, the pressure of a user for cleaning up a memory is reduced, and the user experience is improved. The mobile terminal has the characteristics of strong utilization capability of storage space of the mobile terminal and high humanization degree.
Owner:XUZHOU AISHENG EDUCATION CONSULTING CO LTD

Document processing method and apparatus, and electronic device

The application discloses a document processing method and device and electronic equipment. Relate to the field of artificial intelligence and document processing, the method comprises: carrying out content identification on a target document, obtaining document content, and splitting the document content according to element types to obtain M document elements; obtaining the attribute information and text content of each document element, and determining the association relationship between the M document elements; generating initial graph structure data of the target document according to the association relationship; determining the feature information of each document element according to the attribute information and text content of each document element, and updating the edges in the initial graph structure data according to the feature information to obtain target graph structure data of the target document. Through the application, the problem that the existing document recognition technology in the related art only outputs linear text, resulting in low content acquisition efficiency and low document use efficiency of the document, is solved.
Owner:BEIJING CALORIE INFORMATION TECH CO LTD

A method and system for accurate comparison of Chinese examination and approval documents and actual seal documents

This invention discloses a method and system for accurately comparing Chinese approval documents and actual stamped documents, relating to the field of document recognition technology. The system includes: a first encoding module, a first similarity module, a first judgment module, a second judgment module, a region division module, a second encoding module, a second similarity module, a third judgment module, a character recognition module, a first recognition and judgment module, an image reconstruction module, and a second recognition and judgment module. This invention utilizes multi-region similarity measurement to pre-judge whether Chinese approval documents and actual stamped documents match, and uses the TextBoxes character recognition method and a super-resolution-based accurate recognition method to complete the accurate comparison of Chinese approval documents and actual stamped documents.
Owner:BEIJING WILION TIME TECH

Test case management method and device, equipment and storage medium

The invention relates to the field of software testing, and discloses a test case management method, device and equipment and a storage medium, which can be used for system testing in financial and medical fields, and comprises the following steps: obtaining a current demand document and a current test case corresponding to a current test, the current demand document comprising demand function points required by the current test; identifying missing difference test function points of the current test case according to the current demand document; obtaining a historical test case comprising the difference test function point; and optimizing the current test case according to the historical test case to obtain an optimized test case. The problems of low reuse efficiency of test cases and insufficient coverage of demand function points in the prior art are solved.
Owner:PING AN HEALTH INSURANCE CO LTD

Customs declaration document identification-based customs risk detection method and system

The application provides a customs clearance risk detection method and system based on customs declaration document recognition, and relates to the technical field of customs declaration document risk management. The method comprises the following steps: obtaining a scanned copy of a basic customs declaration document, and determining whether additional customs declaration documents are required; if additional customs declaration documents are required, generating upload prompt information; determining a high customs declaration risk trigger result according to the recognition result of the scanned copy of the basic customs declaration document and the upload result of the scanned copy of the additional customs declaration document; determining a medium customs declaration risk trigger result according to the recognition results of the scanned copy of the basic customs declaration document and the scanned copy of the additional customs declaration document; determining a low customs declaration risk trigger result according to the recognition result of the scanned copy of the basic customs declaration document; and determining a customs clearance risk detection result according to the high customs declaration risk trigger result, the medium customs declaration risk trigger result and the low customs declaration risk trigger result. According to the application, accurate prompts can be given based on the risk level of the customs declaration documents, so as to improve the probability of smooth customs clearance.
Owner:CHINA XIAMEN OCEAN SHIPPING AGENCY CO LTD

Deep learning-based circular packaging logistics document identification and data processing method

The invention relates to the technical field of deep learning and logistics document identification, in particular to a circular packaging logistics document identification and data processing method based on deep learning. The method includes: acquiring an original document image captured on site; obtaining a service context; performing physical state analysis processing on the original document image to obtain a physical state mask and a key information area physical entropy increase index; original data fragments, optical character recognition confidence and model uncertainty are obtained; combining the physical entropy increase index of the key information region, the optical character recognition confidence coefficient and the model uncertainty to solve a coupling data credibility score; based on the coupling data credibility score, the original data fragment and the service context, obtaining credible data; and generating structured data and actionable service suggestions according to credibility scores of the credible data and the coupled data. According to the method, the robustness and the environmental adaptability of circular packaging document identification in a complex field environment are remarkably improved.
Owner:ANWOOD LOGISTICS SYSTEMS (SUZHOU) CO LTD

Method and related device for merging and dividing of OCR document recognition results

ActiveCN117115822Bcoherent contentSorting accurately and smoothlyEnergy efficient computingInstrumentsText recognitionDepth-first search
The application discloses a merging and dividing method for an OCR document recognition result and related devices. The method obtains the text box and the position information corresponding to the text box of the document through an OCR model, then calculates the order and the minimum distance between the text boxes; takes the text box as a node, constructs the directional edge between the nodes according to the order and the minimum distance of the text box, and obtains a directed graph; traverses the directed graph by using a depth-first search algorithm to obtain a plurality of text sets, determines the order between the text sets according to the order of the text box, and obtains a set order; according to the set order, calculates the probability that the two adjacent text sets belong to a sentence above and below, and judges whether the calculation result is greater than a preset probability threshold; if yes, the corresponding two text sets are merged according to the set order to obtain a final text recognition result. Compared with the prior art, the application greatly improves the readability of the text recognition content, and saves time and labor cost.
Owner:WUHAN WANWUYUN DIGITAL OPERATION CO LTD

Coastal bulk cargo paper SOF receipt identification method and system based on large model

The invention discloses a coastal bulk cargo paper SOF document identification method and system based on a large model, and relates to the technical field of image information identification, and the method comprises the steps: taking a paper SOF document image, employing an image detection algorithm of an OCR module to detect a text region of the paper SOF document image, obtaining the coordinate information of each text box, and carrying out the recognition of the coordinate information of each text box; performing direction correction and character recognition on the textbox, and outputting an original recognition result containing text content and corresponding space coordinates; performing joint coding on an original recognition result in the textbox coordinate fusion model, constructing a text association graph by utilizing a geometrical relationship and semantic coherence between textboxes, and generating a structured representation with spatial position features; extracting key fields of a main table based on the structured representation by adopting an SOF extraction model, complementing missing information of the table, extracting time points and discrete information, and outputting final structured data; according to the invention, the information identification efficiency and identification accuracy of the paper SOF document are improved.
Owner:GUANGZHOU ZHENHUA AVIATION TECH CO LTD

Enrollment information structured extraction method and system applied to college entrance examination

The invention discloses an enrollment information structured extraction method and system applied to college entrance examination, and the method comprises the steps: obtaining an enrollment abbreviation of each college, and forming an original document set; for each enrollment brief seal document, segmenting into a plurality of text paragraphs, calculating and comparing the information importance of each paragraph, and sequencing to form a reconstructed document; identifying various types of enrollment rules to be extracted in the reconstructed document, and evaluating the correlation between each text paragraph and the enrollment rules; taking a rule with correlation as focus information, inputting a document subjected to focus injection processing into a large language model, and extracting a structured enrollment rule according to a preset format; and simulating a professional admission process by using the extracted structured enrollment rule in combination with the score information of the examinee, and comparing a simulation result with an actual admission result to generate an audit report. Through importance rearrangement, it is ensured that core enrollment terms can be preferentially and completely processed by the model, and key information is prevented from being ignored due to the fact that the position is close to the back.
Owner:XI AN JIAOTONG UNIV

Multi-format document recognition and enterprise knowledge base construction system and method and storage medium

The invention relates to the technical field of document information processing, and discloses a multi-format document recognition and enterprise knowledge base construction system and method and a storage medium. The system comprises a document collection module used for obtaining a to-be-processed original document set and performing format analysis to obtain a target format unified document set; the metadata complementing module is used for extracting key fields, activating a compensation unit when missing is detected, and generating a metadata complete document set; the semantic division module is used for dividing semantic subunits according to a preset knowledge structure template, and filtering the semantic subunits after feature analysis; the knowledge recombination module is used for extracting feature vectors of the residual semantic subunits, comparing the feature vectors with standard vectors to construct a knowledge unit set, and outputting a result through weight sorting and association verification; and the index construction module is used for generating an index structure according to the verification result and the unassociated vector and updating the enterprise knowledge base. The system can effectively process multi-format documents, improve metadata, mine knowledge association and optimize a knowledge base structure.
Owner:DALIAN POWER SUPPLY COMPANY STATE GRID LIAONING ELECTRIC POWER

Automated document identification and auditing system

An auditing system may process documents associated with a transaction to audit the entire transaction or the documents involved in the transaction. The documents are received and classified by document types. Structured text, unstructured text or both are extracted from the documents and structured data is produced using the extracted text. The documents are audited using the structured data. In some embodiments, the documents may be audited using structured auditing questions, automated programmatic verification, external data sources, or combinations thereof.
Owner:DEALR INC

Information processing device, information processing method, and program

We provide an information processing device that enables sales role-playing that closely resembles actual sales experiences, assuming a target audience with knowledge of the sales target. [Solution] The information processing device 1 comprises an input receiving unit 121 that receives document identification information to identify sales material information, an operation unit 132 that obtains sales material information corresponding to one or more document identification information from one or more document databases 44 in which sales material information corresponding to one or more document identification information is stored, and provides the sales material information and customer persona prompt to the generating AI, an answer acquisition unit 133 that obtains an answer from the generating AI, and an answer output unit 142 that outputs the answer, and further receives training input which is input to advance role-playing, provides the training input to the generating AI, obtains an answer from the generating AI, and outputs the answer.
Owner:AMPTALK CO LTD

Document segmentation method and device, equipment and storage medium

The invention provides a method for segmenting a document. The method comprises the following steps: acquiring a to-be-segmented document; identifying each title in the to-be-segmented document and performing hierarchical identification on each title; and segmenting the to-be-segmented document according to each title and the hierarchical identifier corresponding to each title to obtain each paragraph text. According to the document segmentation method and device, the titles in the to-be-segmented document are recognized, the titles are subjected to hierarchical identification, and then the document is segmented according to the titles with the hierarchical identification, so that the titles in the document can be effectively utilized, and the structure of the document can be accurately reflected through the hierarchical identification of the titles; in this way, the titles in the to-be-segmented document and the structure information contained in the titles can be effectively utilized for segmentation, and the segmentation result is more accurate.
Owner:HANGZHOU FEIZHIYUN INFORMATION TECH CO LTD