Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

40 results about "Document segmentation" patented technology

AI agent construction system and method based on hybrid retrieval and father-child segmentation

The invention discloses an AI (artificial intelligence) agent construction system based on hybrid retrieval and father-child segmentation, which comprises the following steps of: dividing a subclass knowledge base according to domain knowledge, performing father-child segmentation processing, and constructing a hierarchical semantic network; vectorization embedding and deep semantic reconstruction are carried out on the user question text; retrieving the reconstructed problem by adopting a mixed retrieval algorithm combining sparse retrieval and dense retrieval, and forming a high-score sub-segment set according to a comprehensive score obtained by dynamic weight distribution; mapping the sub-segments to the parent segment through a hierarchical backtracking algorithm, aggregating brother nodes to form an extended candidate set, and generating an associated sub-segment set after duplicate removal and re-retrieval; and finally inputting a large language model to generate a complete answer. According to the method, the problems of context segmentation, low retrieval accuracy and complicated knowledge base maintenance of traditional document segments are solved, the answer coverage and accuracy of an intelligent question-answering system are remarkably improved, and the method is suitable for knowledge question-answering scenes in the complicated technical fields such as intelligent network connection automobiles and the like.
Owner:DONGFENG MOTOR GRP

Document review method based on multi-agent cooperation and retrieval enhancement generation

PendingCN121706760ASemantic analysisBiological modelsDocument segmentationGenerative model
The invention discloses a document review method based on multi-agent collaboration and retrieval enhancement generation. The method comprises the following steps: firstly, analyzing a document review rule input by a user into a rule semantic intermediate representation through a natural language processing technology, and calling a retrieval enhancement generation (RAG) module to expand related knowledge to form a structured rule library; secondly, multi-modal analysis and chapter segmentation are carried out on a to-be-examined document, and elements such as texts, tables, pictures and formulas are expressed in a unified mode; tasks such as rule analysis, document segmentation, knowledge retrieval, matching comparison and report generation are completed through a multi-agent cooperation mechanism; finally, when ambiguity exists in rule and document matching, an RAG module is introduced to retrieve supplementary evidences from an external knowledge base, explanatory comparison is conducted in combination with the generative model, and the accuracy and authority of judgment are improved.
Owner:ZHEJIANG UNIV OF TECH

Automatic compliance examination method and system based on graph retrieval enhancement

The invention discloses an automatic compliance examination method and system based on graph retrieval enhancement. The method comprises the following steps of: obtaining a billing specification and a historical inquiry report as input documents; the input document is subjected to dual-mode document segmentation based on a large language model, logic blocks and physical blocks are generated, and each logic block is a continuous page unit and is attached with a content abstract generated by the model; constructing a precedent graph, extracting key legal elements in the historical inquiry report through multi-stage recursion, generating semantic network nodes carrying traceability identifiers, and establishing cross-document semantic association; and constructing a state graph, automatically identifying chapters and logic and semantic relationships among the chapters based on a document hierarchical structure, and generating a machine-readable structured index and the like. According to the method, the long text semantic understanding depth, the cross-section consistency and the legal reasoning accuracy are remarkably improved, the manual review cost is reduced, and the method is suitable for a listing compliance review scene under a registration system.
Owner:AMI INTELLIGENT (XIAMEN) TECHNOLOGY CO LTD

Document segmentation method and device based on large language model, equipment and storage medium

The invention provides a document segmentation method and device based on a large language model, equipment and a storage medium, and relates to the technical field of text processing. The method comprises the steps of inputting a to-be-segmented target document into a pre-trained large language model, and executing the following operations through the large language model: performing text layout analysis on the target document, and identifying titles and all paragraphs of each level in the target document; for each paragraph, inserting an associated title related to the paragraph in all titles into an initial position of the paragraph to obtain a corresponding target paragraph; and sorting all the target paragraphs based on the semantic similarity among all the target paragraphs, and determining a segmentation result of the target document based on all the sorted target paragraphs. By the adoption of the technical scheme, when document segmentation is carried out, semantic loss in the document segmentation process can be effectively reduced, and therefore the document segmentation effect is improved.
Owner:CHINA LIFE ASSET MANAGEMENT CO LTD

Document segmentation method and system based on context marking and model cascading

The invention relates to the technical field of natural language processing, and provides a document segmentation method based on context marking and model cascading, which comprises the following steps: in response to a document segmentation request, loading a to-be-processed document and initializing segmentation parameters; calling a large language model to analyze a current to-be-processed text segment, identifying a logic demarcation point, and generating wedge information containing a segmentation point mark and a context; positioning absolute positions of segmentation points in the text segment according to the wedge information, and segmenting the text segment into a plurality of text sub-segments; repeating the execution until a preset recursion termination condition is reached; after recursion is completed, segmenting results of all layers are aggregated, a hierarchical document structure is constructed, and the segmenting results are output. Through a wedge mechanism, only tiny positioning marks are output, the token cost is reduced, and the overall cost is optimized in combination with a model cascading strategy. And the generated text block is highly aligned with the semantic boundary of the document, so that the context fragmentation problem is effectively solved, the context relevance is enhanced, and the model illusion is inhibited.
Owner:SHANGHAI WENYIN INTERNET INFORMATION TECHNOLOGY CO LTD

Question and answer method and device, electronic equipment and computer storage medium

The invention discloses a question and answer method and device, electronic equipment and a computer storage medium, and the method comprises the steps: carrying out the intention recognition of collected original question information of a user, and obtaining an intention recognition result; and performing vector matching in a preset knowledge base according to the intention recognition result, and searching in the preset knowledge base to obtain knowledge points corresponding to the original question information of the user and pictures matched with the knowledge points. And outputting knowledge points corresponding to the original question information of the user and pictures matched with the knowledge points. Through the multi-modal document segmentation method, the picture can be bound with the previous paragraph or the following paragraph after being recognized in the segmentation process, and the picture can be taken out together as a reference picture when the output text relates to the previous or following content, so that knowledge loss is avoided.
Owner:CHINA MOBILE COMM GRP SHAANXI CO LTD +1

A Fast Document Feature Extraction System Based on Pre-trained Large Models

This invention provides a rapid document element extraction system based on a pre-trained large model, belonging to the field of computer software application technology. The system includes: a parameter domain adaptation module for textualizing documents and constructing an industry-standard corpus based on the textualization results, then adjusting a pre-defined language model using the industry-standard corpus; a dynamic document segmentation module for semantically segmenting industry-standard documents to obtain several text blocks; an entity alignment module for extracting entities and relations from the text blocks and performing entity alignment using uniform manifold approximation and projection methods; and a relational reasoning and knowledge graph completion module for completing a preliminary knowledge graph and storing the completion results. This invention eliminates the need for pre-defined rule templates or data annotation, directly improving element extraction efficiency.
Owner:ANHUI BIAOXINCHA DATA TECH CO LTD

A machine learning system for automatic document segmentation and classification

A computer-implemented method for automatically splitting and classifying an input document into one or more sub-documents using a machine learning system is described. The machine learning system includes a visual segmentation neural network, an optical character recognition subsystem, a title classifier, a document classifier, and a grouper subsystem. The method includes receiving visual input representing a plurality of pages of the input document, using the visual segmentation neural network to classify each page of the input document into each of a plurality of templates, determining, for each page of the input document, the final document type to which the page belongs, and using the grouper subsystem to group the plurality of pages of the input document into one or more sub-documents based on (i) each template of each page and (ii) each final document type to which each page belongs.
Owner:FPT USA CORP

Hierarchical domain ontology and knowledge graph construction method and device

PendingCN122366609AData packSemantic alignment
This application discloses a method, apparatus, and device for constructing a hierarchical domain ontology and knowledge graph. The method includes: acquiring multi-source input data and performing document segmentation processing on the multi-source input data to obtain parsable document fragments. The multi-source input data includes at least one of device service relationships, operating procedures, and system topology. Based on AI, metadata is extracted from the document fragments to construct a metadata hierarchy tree corresponding to multiple semantic levels. Through a semantic alignment engine and adapter module, the introduced open-source external ontology is semantically mapped and aligned with the multiple semantic levels. The aligned metadata and external ontology knowledge are organized, classified and layered according to six semantic levels to form a structured hierarchy tree, and a hierarchical domain ontology and corresponding RDF knowledge graph are generated based on the structured hierarchy tree to improve the support of the knowledge graph for complex mechanism modeling and reasoning, as well as its usability and maintainability.
Owner:PERSAGY TECHNOLOGY CO LTD

Intelligent tendering and bidding question-answering method based on document segmentation

The invention discloses an intelligent tendering and bidding question-answering method based on document segmentation, and relates to the technical field of document segmentation, and the method comprises the steps: carrying out the layout analysis of a tendering and bidding document, so as to extract a structural unit set, construct a structural hierarchical relation, and carrying out the sentence processing of each structural unit; constructing a semantic window for the structural unit and generating a mixed semantic window vector, and performing primary segmentation on the structural unit by calculating a semantic gradient to generate a semantic continuous block set; second-level segmentation is carried out on the basis of the length constraint of the token, a text block set is generated, cross-level semantic coding is carried out on text blocks, and block-level vector codes are generated; outputting a local abstract for each text block according to a cross-level coding result, constructing a local abstract set, performing compression to generate a global abstract, and constructing a bidirectional mapping system of clause numbers and the text blocks; and efficient and traceable bidding and tendering document analysis and intelligent question and answer are provided for the user.
Owner:JIANGSU PROVINCIAL TENDERING CENTER CO LTD

Method and device for matting and displaying magnetic files on whiteboard

The invention provides a method and device for matting and displaying a magnetic file on a whiteboard, and the method comprises the steps: respectively collecting whiteboard data and magnetic file data, and correspondingly training a whiteboard detection model and a magnetic file segmentation model; performing whiteboard positioning and segmentation on an original image containing a magnetic file based on the whiteboard detection model to obtain a whiteboard area image; performing magnetic suction file segmentation on the white board area image based on the magnetic suction file segmentation model to obtain a target magnetic suction file image; and establishing association between the target magnetic file image and the corrected and amplified image through a mapping table, and realizing real-time matting display of the magnetic file. When a speaker demonstrates and explains the magnetic files on the whiteboard, only the effect presented by the magnetic files is achieved on a picture.
Owner:BEIJING MYSHER TECH

Document segmentation methods, apparatus, computer equipment and storage media

ActiveCN121902816BAccurately identify structural boundariesImprove Segmentation AccuracyDocument structuringEngineering
This application discloses a document segmentation method, apparatus, computer device, and storage medium. In response to a segmentation command, a document to be segmented is acquired; the document is parsed to obtain multiple text units; visual features related to the document layout are determined based on the text units; a segmentation score is calculated based on the visual features; and segmentation is performed based on the segmentation score. In this application, the visual layout information of the document is referenced from a visual layout perspective, avoiding the text extraction quality defects of treating the document as a plain text stream without relying on OCR. Instead, segmentation is performed by combining the visual geometric layout characteristics of the document when the user browses the text, conforming to the browsing patterns of users reading documents, accurately identifying document structural boundaries, and improving the accuracy of text segmentation.
Owner:HANGZHOU YOUZAN TECH CO LTD

A multi-page document positioning and printing system, apparatus, medium and product

PendingCN122507700AText recognitionPage (document)
A multi-page document positioning and printing system, apparatus, medium, and product are disclosed, relating to the field of printing systems. In this system, a data mapping module establishes a mapping relationship between first and second encoded information and stores it in a database; a document segmentation module splits the target document into independent units by page; a content recognition module performs intelligent text recognition on each single-page document and extracts key encoded information; a remote triggering module enables the collection of first encoded information via a mobile terminal and remote triggering of background processing, allowing operators to work flexibly in various areas of the warehouse without having to travel to fixed workstations; and a printing execution module accurately locates the single-page document containing the corresponding encoded information based on the established mapping relationship and drives printing. The entire system not only improves the operational efficiency of warehouse packing and shipping processes and reduces the error rate of searching and matching, but also meets the high requirements of cross-border e-commerce warehousing and logistics for rapid response and precise operation.
Owner:GUANGZHOU JIAOYUN YICHENG CLOTHING CO LTD

Visual analysis method, device and equipment for dynamic threshold value of transformer oil chromatography and medium

PendingCN121899317AImplement exception recognitionImprove powerComponent separationElectric power systemEngineering
The embodiment of the invention discloses a document segmentation method and device, computer equipment and a storage medium. In response to a segmentation instruction, the embodiment of the invention provides a transformer oil chromatography dynamic threshold visual analysis method and device, equipment and a medium. A plurality of oil-immersed transformers are arranged in a power system, and the method comprises the following steps: acquiring oil chromatography data, operation condition data and historical data monitored on line, and establishing a dynamic health threshold baseline for evaluating the state of equipment based on the historical data; and generating a corresponding early warning signal and performing fault early warning based on the deviation distance, the deviation direction and the deviation mode of the deviation dynamic health threshold baseline determined by the oil chromatographic data and the operation condition. According to the method, different from static analysis of dissolved gas in oil based on a static threshold value, oil chromatographic data and a dynamic self-adaptive mechanism oriented to operation condition data are fully mined, long-term slowly-changing gas abnormity identification is realized, and the fault early warning precision of a power system is improved.
Owner:MAINTENANCE BRANCH COMPANY STATE GRID ZHEJIANG ELECTRIC POWER

Database construction method and database system based on hybrid index retrieval

PendingCN122346511AData segmentSemantic search
The application relates to a database construction method and a database system based on mixed index retrieval, and the database construction method comprises the following steps: step one, performing multi-granularity segmentation on collected data to form data segments at different levels; step two, marking key information for each data segment; step three, simultaneously establishing a field index and a vector index for each data segment; step four, establishing a mapping relationship among the collected data, the data segments and the key information; and step five, based on a user query task, simultaneously calling the field index and the vector index for retrieval, fusing and sorting preliminary retrieval results, and generating final retrieval results. The database comprises a document segmentation module, a key information marking module, a mixed index construction module, an association mapping establishment module and a retrieval fusion module. The application can simultaneously consider accurate retrieval and semantic retrieval, improves a database hit rate, supports fine-granularity positioning of parameters and evidence, and improves checkability of retrieval results.
Owner:HARBIN INST OF TECH

Semantic-based case document cutting method and device and related equipment

The embodiment of the invention discloses a case document cutting method and device based on semantics and related equipment. The method comprises the following steps: acquiring a case document, and preprocessing the case document to obtain a standard text; semantic recognition is conducted on the standard text according to the standard logic structure, and entities and entity relations in the standard text are recognized; building a semantic tree of the case document based on the entities and the entity relationship; evaluating the element integrity of each sub-tree of the semantic tree based on a pre-trained large language model, and judging whether the element integrity is greater than a preset threshold; if the element integrity of the sub-tree is greater than a threshold value, taking the document content corresponding to the sub-tree as a document fragment to perform document segmentation; if the element integrity of the sub-tree is lower than the threshold value, text length detection is conducted on the sub-tree, and the content of the sub-tree is optimized according to the text length till the element integrity of the sub-tree is larger than the integrity threshold value. According to the method, the damage of the traditional cutting mode to the internal logic of the document is avoided, and the case can be better understood and used.
Owner:BEIJING TAIXIN TIANCHENG TECHNOLOGY CO LTD

A method for extracting information from bidding documents

This application relates to the field of text processing, and more particularly to a method for extracting information from bidding documents. The method includes: segmenting the bidding document into pages, identifying the corresponding text for each page; generating supplementary text descriptions for images and tables within the pages and appending them to the end of the corresponding text to form an enhanced text block sequence; matching tags from the text block sequence according to a pre-built hierarchical tagging system, and generating corresponding prompt word templates based on the tags and a pre-built prompt word template library; inputting the prompt word templates, the enhanced text block sequence, and the contextual text summary as a combined input to a large language model to obtain a structured extraction result with hierarchical relationships; matching the extracted entity content with a local dictionary, and after successful matching, aggregating and organizing the results to output a structured data file. This method reduces the risk of illusions in the generated content without requiring model retraining.
Owner:SHANGHAI MECHANICAL & ELECTRICAL EQUIP TENDERING CO LTD

Document segmentation method and electronic equipment

The invention relates to the technical field of data processing, and discloses a document segmentation method and electronic device.The document segmentation method comprises the steps that a to-be-segmented document is obtained; determining a segmentation mode corresponding to the document according to the file format of the document; segmenting the document according to the segmentation mode of the document to obtain a plurality of text blocks; performing feature extraction on the plurality of text blocks to obtain text vectors; and determining the text vector and the text content corresponding to the text vector as a document segmentation result. According to the method for automatically segmenting the contents of the documents in the diversified formats, the segmentation modes corresponding to the documents are determined according to the file formats of the different documents, then the documents are segmented according to the segmentation modes, the document segmentation results are obtained, and the document content segmentation efficiency can be improved.
Owner:SHENZHEN SHULIAN TIANXIA INTELLIGENT TECH CO LTD

Transnational culture Agent consistency content generation method based on retrieval enhancement

The invention discloses a retrieval enhancement-based transnational culture Agent consistency content generation method, which comprises the following steps of: firstly, processing multi-source text data, constructing a knowledge base capable of supporting culture retrieval, and respectively establishing a sparse index and a dense vector index to support multi-type retrieval; secondly, performing preliminary screening on queries by adopting a hybrid retrieval strategy to obtain candidate documents, generating a hybrid score in a weighted fusion manner, performing refined reordering on a preliminary candidate set by utilizing a cross encoder, then selecting high-confidence document fragments according to target culture, and performing high-confidence document segmentation on the selected high-confidence document fragments; and a cross-culture enhanced context prompt is constructed in combination with a culture constraint template and a soft prompt, so that a specific culture view angle and value preference can be automatically embedded in a large model in a generation process, and culture consistency contents are generated. The technical problem that a traditional large model is difficult to generate contents conforming to different cultural logics and expression habits under the same input is effectively solved, and the cultural logic accuracy and context adaptability of output contents are ensured.
Owner:LIAONING UNIVERSITY

Knowledge base Chinese document segmentation method and management system thereof

The invention discloses a knowledge base Chinese document segmentation method and a management system thereof, and the method comprises the following steps: S1, receiving a segmentation rule parameter set which is set by a user and comprises a segmentation separator, a segmentation maximum length value and a segmentation overlapping proportion; s2, analyzing the target document and extracting full-text content; s3, performing primary segmentation on the full-text content according to the segmentation separators to generate an initial segmentation set; and S4, performing iteration processing on each segment in the initial segment set. According to the method, a Chinese punctuation-driven integrity guarantee mechanism is created for the first time, full-angle full stop, question marks, exclamation marks, province marks and score marks are accurately positioned as sentence boundaries, and it is ensured that the head and the tail of each segment are complete Chinese sentences; designing an overlap transfer algorithm, extracting semantic segments containing Chinese punctuations from the tail of the segment, splicing the semantic segments to the head of the next segment, and constructing a context bridge between the segments; a pure Chinese punctuation adaptation system thoroughly eliminates segmentation errors caused by mixed use of Chinese and English.
Owner:GUANGDONG CHENGZHI TECH CO LTD

Electric power part image enhancement and diffusion system for unmanned aerial vehicle electric power inspection

PendingCN122049108AImage enhancementImage analysisEngineeringDocument segmentation
The invention discloses an electric power component image enhancement diffusion system for unmanned aerial vehicle electric power inspection. The system comprises a JSON file, an SAM segmentation and mask sampling module, a color statistics and JSON storage module, a CLIP semantic matching module, an operator screening module and a Stable Diffusion diffusion module. Component fields, dynamic data and a text description library are preset in the JSON file; the SAM segmentation and mask sampling module takes an original image of a power component and an SAM2 model as input and output PNG format enhanced mask files and CSV format color data tables; the color statistics and JSON storage module determines a main color and an auxiliary color by using a K-means clustering algorithm and stores the main color and the auxiliary color in a JSON file; the CLIP semantic matching module is used for calculating cosine similarity of visual features and text features and screening an optimal operator; the operator screening module splits basic operators and outputs positive and negative Prompt lists corresponding to final control calculations; and the Stable Diffusion diffusion module takes the original image of the power component and the enhanced mask file as input to generate an enhanced image of the power component.
Owner:TIANJIN UNIV

A method and system for detecting a power transmission line fitting by using a SAM model to amplify a sample

ActiveCN117253159BHigh quality training dataAchieve expansionCharacter and pattern recognitionNeural learning methodsSmall sampleAlgorithm
This invention relates to a method and system for detecting transmission line fittings using SAM model augmentation samples. The method includes: sorting images of the types of fittings to be detected based on UAV imagery; labeling the fittings in the images; segmenting the fitting samples from the images according to the corresponding labeling files; pasting the data-augmented fitting samples into the images to construct a training dataset with a balanced distribution of fitting sample sizes; and iteratively training the target detection model using the training data to obtain a high-precision multi-category fitting detection model. This invention combines multiple data augmentation techniques to generate fitting samples, amplifying the original small sample sizes of fittings, and ultimately generating a data-trained target detection model with a balanced distribution of fitting sample sizes, achieving high-precision multi-category fitting detection.
Owner:YUNNAN ELECTRIC POWER TESTING & RES INST (GRP) CO LTD

Document segmentation method and device, equipment and storage medium

The invention provides a method for segmenting a document. The method comprises the following steps: acquiring a to-be-segmented document; identifying each title in the to-be-segmented document and performing hierarchical identification on each title; and segmenting the to-be-segmented document according to each title and the hierarchical identifier corresponding to each title to obtain each paragraph text. According to the document segmentation method and device, the titles in the to-be-segmented document are recognized, the titles are subjected to hierarchical identification, and then the document is segmented according to the titles with the hierarchical identification, so that the titles in the document can be effectively utilized, and the structure of the document can be accurately reflected through the hierarchical identification of the titles; in this way, the titles in the to-be-segmented document and the structure information contained in the titles can be effectively utilized for segmentation, and the segmentation result is more accurate.
Owner:HANGZHOU FEIZHIYUN INFORMATION TECH CO LTD

Document segmentation method and device, computer equipment and storage medium

The embodiment of the invention discloses a document segmentation method and device, computer equipment and a storage medium. In response to the segmentation instruction, obtaining a to-be-segmented document; analyzing the to-be-segmented document to obtain a plurality of text units; determining document visual features related to the document layout based on the text unit; calculating a segmentation score based on the visual features of the document; and performing segmentation based on the segmentation score. In the application, the visual layout information of the document is referred to from the visual layout angle, and the document is no longer used as a plain text stream and does not depend on the text extraction quality defect of OCR (Optical Character Recognition). According to the method, the document is segmented in combination with the visual geometric layout characteristics of the document when the user browses the text, the browsing rule of the user for reading the document is met, the document structure boundary is accurately recognized, and the segmentation precision of the text is improved.
Owner:HANGZHOU YOUZAN TECH CO LTD

Document segmentation methods, devices, electronic devices, storage media, and program products

This invention provides a document segmentation method, apparatus, electronic device, storage medium, and program product, relating to the field of document processing technology. The method includes: acquiring a converted image of the current page in a document to be segmented, and extracting multiple local visual features from the current page image; performing attention fusion calculation based on the multiple local visual features and the visual feature set of the preceding page to generate multiple fused visual features containing global contextual information; inputting the multiple fused visual features and the semantic segmentation result of the preceding page into a semantic segmentation model to obtain the semantic segmentation result of the current page to be processed output by the semantic segmentation model; and segmenting the document to be segmented based on the semantic segmentation results of each page in the document to be segmented to obtain multiple segmented blocks of the document. This invention can utilize the global visual contextual information of the entire document when processing any single page, thereby improving the accuracy of document segmentation.
Owner:IFLYTEK CO LTD

Natural language based document processing method, apparatus and computer program product

The application provides a natural language-based document processing method, comprising the following steps: identifying the type of an original document by using a pre-trained document recognition model; processing text segments of the original document based on natural language technology; quickly matching the belonging chapter of the text segments by rule matching; in the case of rule matching failure, loading a corresponding chapter model according to the type of the document, and identifying the belonging chapter of the text segments based on the chapter model; and generating a structured chapter list according to the chapter belonging results of all text segments of the original document. The model structure of the application is easy to maintain and extend, the document segmentation and information extraction performance is good, and the application is especially suitable for enterprise-level document processing, and can effectively improve the work efficiency of employees.
Owner:DALIAN YUNLU TECHNOLOGY CO LTD

Test data processing method and related equipment

The invention discloses a test data processing method and related equipment, and relates to the technical field of vehicle testing, and the method comprises the steps: carrying out the video recording based on a target vehicle-mounted camera corresponding to a preset test task, and obtaining video cache data; performing file segmentation processing on the video cache data to obtain a plurality of video segmentation files; uploading the plurality of video segment files to a cloud to obtain a plurality of video access links corresponding to the plurality of video segment files; and carrying out association processing on the plurality of video access links and the multi-dimensional test data corresponding to the preset test task to form a test data set corresponding to the preset task. According to the method, the vehicle-mounted video is continuously recorded and managed in a segmented manner, the video is uploaded to the cloud in a segmented manner, the access link is generated, and the access link is associated with the multi-dimensional test data, so that complete acquisition, stable storage, convenient access and unified management of the data in the test process are realized, and the data continuity, traceability and analysis efficiency of the test task can be improved.
Owner:VOYAH AUTOMOBILE TECH CO LTD

Document segmentation method and device, electronic equipment and readable storage medium

The invention discloses a document segmentation method and device, electronic equipment and a readable storage medium, and the method comprises the steps that a semantic structure unit in a to-be-segmented document is recognized, and the semantic structure unit comprises at least one of a causal relationship chain, a process step chain and an entity-action-result mode; the to-be-segmented document is segmented on the basis of the semantic structure unit, at least one first candidate block is obtained, and the at least one first candidate block comprises the first candidate block corresponding to the semantic structure unit; for each first candidate block, under the condition that the length of the first candidate block is greater than a first threshold value, segmenting the first candidate block according to first semantics to obtain at least one target block; for each target block, the target block is associated with target metadata corresponding to the target block, and the target metadata comprises the document position where the target block is located and the semantic structure unit to which the target block belongs.
Owner:BEIJING PERCENT INFORMATION TECH CO LTD

Document segmentation method and device, electronic equipment and readable storage medium

PendingCN122047220Aimprove accuracyAvoid the problem of semantic fragmentationSemantic analysisEngineeringDocumentation
The invention provides a document segmentation method and device, electronic equipment and a readable storage medium. The method comprises the steps that information of a document is subjected to block processing to determine multiple pieces of first block information and first position information of the first block information in the document, and the first block information comprises corresponding first document information in the document; determining directory entry information in the document, wherein the directory entry information comprises second position information of second document information under the corresponding directory entry in the document; according to first position information corresponding to each piece of first block information and second position information corresponding to the directory entry information, determining second block information corresponding to the directory entry information in the plurality of pieces of first block information; and determining segmented fragments of the document based on the second block information corresponding to the directory entry information. According to the invention, the accuracy of document segmentation can be improved.
Owner:RUIJIE NETWORKS CO LTD

Document segmentation method and device, computer equipment and readable storage medium

The invention relates to a document segmentation method and device, computer equipment and a readable storage medium. The method comprises the following steps: acquiring a to-be-segmented document, and calling a target apprenant model; the target apprentice model is obtained by training a sample segmentation result generated based on the prophet model; the prophet model is used for performing automatic fragmentation on a sample document based on semantic mutation point detection to obtain a sample segmentation result serving as a reference; identifying a target semantic mutation point in the to-be-segmented document based on the target apprentice model, and segmenting the to-be-segmented document according to the target semantic mutation point to obtain each initial segmentation result; and determining each segmentation result of the to-be-segmented document according to each initial segmentation result. By adopting the method, the accuracy of the document segmentation method can be improved.
Owner:BEIJING PACTERA JINXIN TECH LTD