Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

60 results about "Document segmentation" patented technology

AI agent construction system and method based on hybrid retrieval and father-child segmentation

The invention discloses an AI (artificial intelligence) agent construction system based on hybrid retrieval and father-child segmentation, which comprises the following steps of: dividing a subclass knowledge base according to domain knowledge, performing father-child segmentation processing, and constructing a hierarchical semantic network; vectorization embedding and deep semantic reconstruction are carried out on the user question text; retrieving the reconstructed problem by adopting a mixed retrieval algorithm combining sparse retrieval and dense retrieval, and forming a high-score sub-segment set according to a comprehensive score obtained by dynamic weight distribution; mapping the sub-segments to the parent segment through a hierarchical backtracking algorithm, aggregating brother nodes to form an extended candidate set, and generating an associated sub-segment set after duplicate removal and re-retrieval; and finally inputting a large language model to generate a complete answer. According to the method, the problems of context segmentation, low retrieval accuracy and complicated knowledge base maintenance of traditional document segments are solved, the answer coverage and accuracy of an intelligent question-answering system are remarkably improved, and the method is suitable for knowledge question-answering scenes in the complicated technical fields such as intelligent network connection automobiles and the like.
Owner:DONGFENG MOTOR GRP

Bidding document information extraction method

The invention relates to the field of text processing, in particular to a bidding document information extraction method. Comprising the following steps: segmenting a bidding and tendering file into pages, and identifying the pages to obtain corresponding texts; generating complementary text description for images and tables in the page and adding the complementary text description to the tail of a text corresponding to the page to form an enhanced text block sequence; matching a label from the text block sequence according to a pre-constructed hierarchical label system, and generating a corresponding cue word template according to the label and a pre-constructed cue word template library; inputting the cue word template, the enhanced text block sequence and the context text abstract as a combination into a large language model to obtain a structured extraction result with a hierarchical relationship; and matching the extracted entity content with a local dictionary, carrying out aggregation arrangement on a result after the matching is passed, and outputting a structured data file. On the premise that the model does not need to be retrained, the illusion risk of the generated content is reduced.
Owner:SHANGHAI MECHANICAL & ELECTRICAL EQUIP TENDERING CO LTD

Document review method based on multi-agent cooperation and retrieval enhancement generation

The invention discloses a document review method based on multi-agent collaboration and retrieval enhancement generation. The method comprises the following steps: firstly, analyzing a document review rule input by a user into a rule semantic intermediate representation through a natural language processing technology, and calling a retrieval enhancement generation (RAG) module to expand related knowledge to form a structured rule library; secondly, multi-modal analysis and chapter segmentation are carried out on a to-be-examined document, and elements such as texts, tables, pictures and formulas are expressed in a unified mode; tasks such as rule analysis, document segmentation, knowledge retrieval, matching comparison and report generation are completed through a multi-agent cooperation mechanism; finally, when ambiguity exists in rule and document matching, an RAG module is introduced to retrieve supplementary evidences from an external knowledge base, explanatory comparison is conducted in combination with the generative model, and the accuracy and authority of judgment are improved.
Owner:ZHEJIANG UNIV OF TECH

Intelligent question answering system implementation method and system

The invention relates to the technical field of intelligent questioning and answering, in particular to an intelligent questioning and answering system implementation method and system.The intelligent questioning and answering system implementation method comprises the following steps of document collection and preprocessing, document dicing, document layout analysis and table layout analysis; the method has the beneficial effects that a semantic association network and a multi-modal index system are formed through offline document collection, preprocessing, slicing, layout analysis, data extraction and knowledge graph construction; in the online part, the capabilities of Embedding vectorization, multi-index joint retrieval, tensor reordering, AI database integration and large language model generation are combined, accurate semantic understanding and rapid knowledge matching of user questions are realized, high-quality answers are generated, and interactive feedback optimization is supported.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

Automatic compliance examination method and system based on graph retrieval enhancement

The invention discloses an automatic compliance examination method and system based on graph retrieval enhancement. The method comprises the following steps of: obtaining a billing specification and a historical inquiry report as input documents; the input document is subjected to dual-mode document segmentation based on a large language model, logic blocks and physical blocks are generated, and each logic block is a continuous page unit and is attached with a content abstract generated by the model; constructing a precedent graph, extracting key legal elements in the historical inquiry report through multi-stage recursion, generating semantic network nodes carrying traceability identifiers, and establishing cross-document semantic association; and constructing a state graph, automatically identifying chapters and logic and semantic relationships among the chapters based on a document hierarchical structure, and generating a machine-readable structured index and the like. According to the method, the long text semantic understanding depth, the cross-section consistency and the legal reasoning accuracy are remarkably improved, the manual review cost is reduced, and the method is suitable for a listing compliance review scene under a registration system.
Owner:AMI INTELLIGENT (XIAMEN) TECHNOLOGY CO LTD

Intelligent self-adaptive document segmentation method oriented to RAG system

The invention relates to the technical field of natural language processing (NLP), in particular to an intelligent self-adaptive document segmentation method oriented to an RAG system. The method comprises the steps that S1, a deep learning model is adopted to dynamically adjust the size and the step length of a window according to document content density, structure information and context semantics; s2, in combination with a window representation result, calculating the context semantic similarity of the segmentation blocks by adopting a language model, and automatically adjusting the size of an overlapping region based on the context semantic similarity; and S3, based on a context segmentation block representation result, associating the segmentation block with the context by introducing a BERT model, automatically adjusting an overlapping part and a window size, and optimizing a segmentation effect. The invention aims to provide the intelligent self-adaptive document segmentation method oriented to the RAG system so as to improve the document segmentation efficiency and quality and optimize the recall rate and the generation quality of the RAG system.
Owner:FUJIAN YIRONG INFORMATION TECH

Multistage document segmentation method adopting self-adaptive dynamic partitioning algorithm

The invention discloses a multi-level document segmentation method adopting an adaptive dynamic partitioning algorithm, and relates to the technical field of document segmentation, the method comprises the following steps: determining a target document type and target chapter information of a to-be-segmented document; when the target document type is a standard document, calculating a target information density corresponding to the to-be-segmented document based on the target chapter information; and based on the target information density, determining a target chapter overlapping degree corresponding to the to-be-segmented document, and segmenting each target chapter according to the target chapter overlapping degree to complete segmentation processing of the to-be-segmented document. According to the method and device, it is ensured that the chapter content of the segmented document is complete and logically coherent, the problems of information loss or chapter content segmentation and the like caused by blind segmentation are solved, the document segmentation quality and efficiency are effectively improved, and the segmented document better meets the actual use requirement.
Owner:DIGITAL HEALTH CHINA TECHNOLOGIES CO LTD

Document segmentation method and device based on large language model, equipment and storage medium

The invention provides a document segmentation method and device based on a large language model, equipment and a storage medium, and relates to the technical field of text processing. The method comprises the steps of inputting a to-be-segmented target document into a pre-trained large language model, and executing the following operations through the large language model: performing text layout analysis on the target document, and identifying titles and all paragraphs of each level in the target document; for each paragraph, inserting an associated title related to the paragraph in all titles into an initial position of the paragraph to obtain a corresponding target paragraph; and sorting all the target paragraphs based on the semantic similarity among all the target paragraphs, and determining a segmentation result of the target document based on all the sorted target paragraphs. By the adoption of the technical scheme, when document segmentation is carried out, semantic loss in the document segmentation process can be effectively reduced, and therefore the document segmentation effect is improved.
Owner:CHINA LIFE ASSET MANAGEMENT CO LTD

Document segmentation method and system based on context marking and model cascading

The invention relates to the technical field of natural language processing, and provides a document segmentation method based on context marking and model cascading, which comprises the following steps: in response to a document segmentation request, loading a to-be-processed document and initializing segmentation parameters; calling a large language model to analyze a current to-be-processed text segment, identifying a logic demarcation point, and generating wedge information containing a segmentation point mark and a context; positioning absolute positions of segmentation points in the text segment according to the wedge information, and segmenting the text segment into a plurality of text sub-segments; repeating the execution until a preset recursion termination condition is reached; after recursion is completed, segmenting results of all layers are aggregated, a hierarchical document structure is constructed, and the segmenting results are output. Through a wedge mechanism, only tiny positioning marks are output, the token cost is reduced, and the overall cost is optimized in combination with a model cascading strategy. And the generated text block is highly aligned with the semantic boundary of the document, so that the context fragmentation problem is effectively solved, the context relevance is enhanced, and the model illusion is inhibited.
Owner:SHANGHAI WENYIN INTERNET INFORMATION TECHNOLOGY CO LTD

Question and answer method and device, electronic equipment and computer storage medium

The invention discloses a question and answer method and device, electronic equipment and a computer storage medium, and the method comprises the steps: carrying out the intention recognition of collected original question information of a user, and obtaining an intention recognition result; and performing vector matching in a preset knowledge base according to the intention recognition result, and searching in the preset knowledge base to obtain knowledge points corresponding to the original question information of the user and pictures matched with the knowledge points. And outputting knowledge points corresponding to the original question information of the user and pictures matched with the knowledge points. Through the multi-modal document segmentation method, the picture can be bound with the previous paragraph or the following paragraph after being recognized in the segmentation process, and the picture can be taken out together as a reference picture when the output text relates to the previous or following content, so that knowledge loss is avoided.
Owner:CHINA MOBILE COMM GRP SHAANXI CO LTD +1

Audio enhancement of video through video file segmentation, event extraction, and contextual data structuring forefficient matching, generation, and / or alignment of audio to adepicted event

Disclosed are a method, a device, and / or a system of audio enhancement of video through video file segmentation, event extraction, and contextual data structuring for efficient matching, generation, and / or alignment of audio to a depicted event. In one embodiment, a system includes a memory storing computer readable instructions that when executed initiate a video object in a database representing a video file and store a video segmentation reference drawn from the video object to a segmentation object, which may represent a shot or scene in the video. The system may parse the video file to extract an event including an event range, an event description, and an event ontology, and may generate encoding vector(s) therefrom. The system may initiate an event object, then link the event object to the video object through the segmentation object, to enable efficient import of context for audio matching and / or audio generation for the event.
Owner:NOCTAL INC

Full-text retrieval method and system, computer equipment and computer readable storage medium

The invention belongs to the field of information retrieval, particularly relates to a full-text retrieval method and system, computer equipment and a computer readable storage medium, and aims to solve the problem of improving the full-text retrieval accuracy. The method comprises the following steps: segmenting a document entering a corpus; calculating paragraph weights of words contained in each segmented document in the segmented document; calculating the document weight of the word in the document according to the paragraph weight of the word; calculating the query weight of the query word, wherein the calculation method of the query weight is the same as the calculation method of the paragraph weight; determining a corresponding target word in a corpus according to the query word; respectively calculating one or more query relevancy according to the query weight and the document weight of the target word; and taking the document corresponding to the document weight corresponding to the maximum n query relevancy as a query result. Semantic features are introduced into weight calculation, and document segmentation processing is combined, so that the full-text retrieval accuracy is effectively improved.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD +1

A Fast Document Feature Extraction System Based on Pre-trained Large Models

This invention provides a rapid document element extraction system based on a pre-trained large model, belonging to the field of computer software application technology. The system includes: a parameter domain adaptation module for textualizing documents and constructing an industry-standard corpus based on the textualization results, then adjusting a pre-defined language model using the industry-standard corpus; a dynamic document segmentation module for semantically segmenting industry-standard documents to obtain several text blocks; an entity alignment module for extracting entities and relations from the text blocks and performing entity alignment using uniform manifold approximation and projection methods; and a relational reasoning and knowledge graph completion module for completing a preliminary knowledge graph and storing the completion results. This invention eliminates the need for pre-defined rule templates or data annotation, directly improving element extraction efficiency.
Owner:ANHUI BIAOXINCHA DATA TECH CO LTD

A machine learning system for automatic document segmentation and classification

A computer-implemented method for automatically splitting and classifying an input document into one or more sub-documents using a machine learning system is described. The machine learning system includes a visual segmentation neural network, an optical character recognition subsystem, a title classifier, a document classifier, and a grouper subsystem. The method includes receiving visual input representing a plurality of pages of the input document, using the visual segmentation neural network to classify each page of the input document into each of a plurality of templates, determining, for each page of the input document, the final document type to which the page belongs, and using the grouper subsystem to group the plurality of pages of the input document into one or more sub-documents based on (i) each template of each page and (ii) each final document type to which each page belongs.
Owner:FPT USA CORP

Hierarchical domain ontology and knowledge graph construction method and device

PendingCN122366609AData packSemantic alignment
This application discloses a method, apparatus, and device for constructing a hierarchical domain ontology and knowledge graph. The method includes: acquiring multi-source input data and performing document segmentation processing on the multi-source input data to obtain parsable document fragments. The multi-source input data includes at least one of device service relationships, operating procedures, and system topology. Based on AI, metadata is extracted from the document fragments to construct a metadata hierarchy tree corresponding to multiple semantic levels. Through a semantic alignment engine and adapter module, the introduced open-source external ontology is semantically mapped and aligned with the multiple semantic levels. The aligned metadata and external ontology knowledge are organized, classified and layered according to six semantic levels to form a structured hierarchy tree, and a hierarchical domain ontology and corresponding RDF knowledge graph are generated based on the structured hierarchy tree to improve the support of the knowledge graph for complex mechanism modeling and reasoning, as well as its usability and maintainability.
Owner:PERSAGY TECHNOLOGY CO LTD

Intelligent tendering and bidding question-answering method based on document segmentation

The invention discloses an intelligent tendering and bidding question-answering method based on document segmentation, and relates to the technical field of document segmentation, and the method comprises the steps: carrying out the layout analysis of a tendering and bidding document, so as to extract a structural unit set, construct a structural hierarchical relation, and carrying out the sentence processing of each structural unit; constructing a semantic window for the structural unit and generating a mixed semantic window vector, and performing primary segmentation on the structural unit by calculating a semantic gradient to generate a semantic continuous block set; second-level segmentation is carried out on the basis of the length constraint of the token, a text block set is generated, cross-level semantic coding is carried out on text blocks, and block-level vector codes are generated; outputting a local abstract for each text block according to a cross-level coding result, constructing a local abstract set, performing compression to generate a global abstract, and constructing a bidirectional mapping system of clause numbers and the text blocks; and efficient and traceable bidding and tendering document analysis and intelligent question and answer are provided for the user.
Owner:JIANGSU PROVINCIAL TENDERING CENTER CO LTD

Method and device for matting and displaying magnetic files on whiteboard

The invention provides a method and device for matting and displaying a magnetic file on a whiteboard, and the method comprises the steps: respectively collecting whiteboard data and magnetic file data, and correspondingly training a whiteboard detection model and a magnetic file segmentation model; performing whiteboard positioning and segmentation on an original image containing a magnetic file based on the whiteboard detection model to obtain a whiteboard area image; performing magnetic suction file segmentation on the white board area image based on the magnetic suction file segmentation model to obtain a target magnetic suction file image; and establishing association between the target magnetic file image and the corrected and amplified image through a mapping table, and realizing real-time matting display of the magnetic file. When a speaker demonstrates and explains the magnetic files on the whiteboard, only the effect presented by the magnetic files is achieved on a picture.
Owner:BEIJING MYSHER TECH

Document segmentation methods, apparatus, computer equipment and storage media

ActiveCN121902816BAccurately identify structural boundariesImprove Segmentation AccuracyDocument structuringEngineering
This application discloses a document segmentation method, apparatus, computer device, and storage medium. In response to a segmentation command, a document to be segmented is acquired; the document is parsed to obtain multiple text units; visual features related to the document layout are determined based on the text units; a segmentation score is calculated based on the visual features; and segmentation is performed based on the segmentation score. In this application, the visual layout information of the document is referenced from a visual layout perspective, avoiding the text extraction quality defects of treating the document as a plain text stream without relying on OCR. Instead, segmentation is performed by combining the visual geometric layout characteristics of the document when the user browses the text, conforming to the browsing patterns of users reading documents, accurately identifying document structural boundaries, and improving the accuracy of text segmentation.
Owner:HANGZHOU YOUZAN TECH CO LTD

A multi-page document positioning and printing system, apparatus, medium and product

PendingCN122507700AText recognitionPage (document)
A multi-page document positioning and printing system, apparatus, medium, and product are disclosed, relating to the field of printing systems. In this system, a data mapping module establishes a mapping relationship between first and second encoded information and stores it in a database; a document segmentation module splits the target document into independent units by page; a content recognition module performs intelligent text recognition on each single-page document and extracts key encoded information; a remote triggering module enables the collection of first encoded information via a mobile terminal and remote triggering of background processing, allowing operators to work flexibly in various areas of the warehouse without having to travel to fixed workstations; and a printing execution module accurately locates the single-page document containing the corresponding encoded information based on the established mapping relationship and drives printing. The entire system not only improves the operational efficiency of warehouse packing and shipping processes and reduces the error rate of searching and matching, but also meets the high requirements of cross-border e-commerce warehousing and logistics for rapid response and precise operation.
Owner:GUANGZHOU JIAOYUN YICHENG CLOTHING CO LTD

Sample set construction method of power grid data and application of sample set construction method in power grid management

The embodiment of the invention provides a sample set construction method of power grid data and application of the sample set construction method in power grid management, and belongs to the technical field of data processing. The sample set construction method comprises the following steps: acquiring fault data scheduled by a distribution network in a power grid as sample data and importing the sample data into a data set; labeling the sample data imported into the data set; performing classification management on the labeled sample data; performing duplicate removal processing on different types of sample data in the data set; performing document segmentation on the sample data in the data set after the duplicate removal processing; and performing knowledge feature extraction on sample data in the data set after document segmentation, and storing the sample data in a database to obtain a sample set. The sample set construction method can collect and classify power grid data.
Owner:XUANCHENG POWER SUPPLY OF ANHUI ELECTRIC POWER CORP +1

Visual analysis method, device and equipment for dynamic threshold value of transformer oil chromatography and medium

PendingCN121899317AImplement exception recognitionImprove powerComponent separationElectric power systemEngineering
The embodiment of the invention discloses a document segmentation method and device, computer equipment and a storage medium. In response to a segmentation instruction, the embodiment of the invention provides a transformer oil chromatography dynamic threshold visual analysis method and device, equipment and a medium. A plurality of oil-immersed transformers are arranged in a power system, and the method comprises the following steps: acquiring oil chromatography data, operation condition data and historical data monitored on line, and establishing a dynamic health threshold baseline for evaluating the state of equipment based on the historical data; and generating a corresponding early warning signal and performing fault early warning based on the deviation distance, the deviation direction and the deviation mode of the deviation dynamic health threshold baseline determined by the oil chromatographic data and the operation condition. According to the method, different from static analysis of dissolved gas in oil based on a static threshold value, oil chromatographic data and a dynamic self-adaptive mechanism oriented to operation condition data are fully mined, long-term slowly-changing gas abnormity identification is realized, and the fault early warning precision of a power system is improved.
Owner:MAINTENANCE BRANCH COMPANY STATE GRID ZHEJIANG ELECTRIC POWER

Database construction method and database system based on hybrid index retrieval

PendingCN122346511AData segmentSemantic search
The application relates to a database construction method and a database system based on mixed index retrieval, and the database construction method comprises the following steps: step one, performing multi-granularity segmentation on collected data to form data segments at different levels; step two, marking key information for each data segment; step three, simultaneously establishing a field index and a vector index for each data segment; step four, establishing a mapping relationship among the collected data, the data segments and the key information; and step five, based on a user query task, simultaneously calling the field index and the vector index for retrieval, fusing and sorting preliminary retrieval results, and generating final retrieval results. The database comprises a document segmentation module, a key information marking module, a mixed index construction module, an association mapping establishment module and a retrieval fusion module. The application can simultaneously consider accurate retrieval and semantic retrieval, improves a database hit rate, supports fine-granularity positioning of parameters and evidence, and improves checkability of retrieval results.
Owner:HARBIN INST OF TECH

Semantic-based case document cutting method and device and related equipment

The embodiment of the invention discloses a case document cutting method and device based on semantics and related equipment. The method comprises the following steps: acquiring a case document, and preprocessing the case document to obtain a standard text; semantic recognition is conducted on the standard text according to the standard logic structure, and entities and entity relations in the standard text are recognized; building a semantic tree of the case document based on the entities and the entity relationship; evaluating the element integrity of each sub-tree of the semantic tree based on a pre-trained large language model, and judging whether the element integrity is greater than a preset threshold; if the element integrity of the sub-tree is greater than a threshold value, taking the document content corresponding to the sub-tree as a document fragment to perform document segmentation; if the element integrity of the sub-tree is lower than the threshold value, text length detection is conducted on the sub-tree, and the content of the sub-tree is optimized according to the text length till the element integrity of the sub-tree is larger than the integrity threshold value. According to the method, the damage of the traditional cutting mode to the internal logic of the document is avoided, and the case can be better understood and used.
Owner:BEIJING TAIXIN TIANCHENG TECHNOLOGY CO LTD

A method for extracting information from bidding documents

This application relates to the field of text processing, and more particularly to a method for extracting information from bidding documents. The method includes: segmenting the bidding document into pages, identifying the corresponding text for each page; generating supplementary text descriptions for images and tables within the pages and appending them to the end of the corresponding text to form an enhanced text block sequence; matching tags from the text block sequence according to a pre-built hierarchical tagging system, and generating corresponding prompt word templates based on the tags and a pre-built prompt word template library; inputting the prompt word templates, the enhanced text block sequence, and the contextual text summary as a combined input to a large language model to obtain a structured extraction result with hierarchical relationships; matching the extracted entity content with a local dictionary, and after successful matching, aggregating and organizing the results to output a structured data file. This method reduces the risk of illusions in the generated content without requiring model retraining.
Owner:SHANGHAI MECHANICAL & ELECTRICAL EQUIP TENDERING CO LTD

A document segmentation method of adaptive slice of large model retrieval enhanced generation

The application discloses a document segmentation method for adaptive segmentation of large model retrieval enhanced generation, and relates to the technical field of large model retrieval enhanced generation, which comprises the following steps: obtaining a document to be segmented, and segmenting the document to be segmented according to a title type to obtain at least one group of original cut blocks; calculating an optimal segmentation number of any original cut block according to information density and theme variation degree corresponding to the original cut block; and segmenting the original cut block according to the optimal segmentation number. According to the application, the document is first segmented according to hierarchical titles, then the information density and the theme variation degree under the hierarchical titles are calculated, the optimal segmentation size under the hierarchical titles is automatically calculated in the unit of the hierarchical titles, and the adaptive segmentation of the document is guided, so that the effect of subsequent retrieval and generation tasks is improved.
Owner:DIGITAL HEALTH CHINA TECHNOLOGIES CO LTD

A method for evaluating document segmentation effects in large-scale model retrieval enhancement generation

The present invention discloses a method for evaluating the effect of document segmentation in large-scale model retrieval enhancement generation, which relates to the technical field of document segmentation. The method includes: obtaining segmentation pairs obtained after segmentation processing of a document to be evaluated, inputting the segmentation pairs into a general semantic model in order, obtaining a target evaluation value corresponding to each segmentation pair, and determining the target effect level corresponding to all target evaluation values ​​based on the corresponding relationship between the evaluation value and the effect level; the training process of the general semantic model is specifically as follows: segmenting the training document to obtain at least two original segments; randomly segmenting any original segment to obtain a preset number of slices; labeling according to whether there is semantic relevance to obtain n groups of training samples; calculating the target relevance score and target separability score corresponding to any group of training samples, and determining the evaluation value corresponding to the training sample. The present invention can feedback the score of the segmentation effect corresponding to each document and can also help assist in document segmentation.
Owner:DIGITAL HEALTH CHINA TECHNOLOGIES CO LTD

Document segmentation method and electronic equipment

The invention relates to the technical field of data processing, and discloses a document segmentation method and electronic device.The document segmentation method comprises the steps that a to-be-segmented document is obtained; determining a segmentation mode corresponding to the document according to the file format of the document; segmenting the document according to the segmentation mode of the document to obtain a plurality of text blocks; performing feature extraction on the plurality of text blocks to obtain text vectors; and determining the text vector and the text content corresponding to the text vector as a document segmentation result. According to the method for automatically segmenting the contents of the documents in the diversified formats, the segmentation modes corresponding to the documents are determined according to the file formats of the different documents, then the documents are segmented according to the segmentation modes, the document segmentation results are obtained, and the document content segmentation efficiency can be improved.
Owner:SHENZHEN SHULIAN TIANXIA INTELLIGENT TECH CO LTD

Rich media intelligent information editing method and device

The embodiment of the application provides a kind of rich media intelligent information editing method and device, method includes: receiving and storing the rich media resource file uploaded by user, the file identification is carried out to the rich media resource file, and according to the file type, file content feature and preset size threshold obtained by identification, the file is segmented by adaptive segmentation algorithm to the rich media resource file, obtains one or more file segments;The file segment is vectorized by preset rich media vectorization model to multi-modal content, and the corresponding semantic meaning array is extracted, the similarity of other rich media resource files and the current rich media resource file is determined according to the semantic meaning array by multidimensional weighted similarity algorithm, and the information of similar rich media resource file with the similarity of current rich media resource file exceeding preset threshold is shown in preset user file editing interface;The application can more flexibly process rich media information.
Owner:BEIJING DUBEN INFORMATION TECHNOLOGY CO LTD

Transnational culture Agent consistency content generation method based on retrieval enhancement

The invention discloses a retrieval enhancement-based transnational culture Agent consistency content generation method, which comprises the following steps of: firstly, processing multi-source text data, constructing a knowledge base capable of supporting culture retrieval, and respectively establishing a sparse index and a dense vector index to support multi-type retrieval; secondly, performing preliminary screening on queries by adopting a hybrid retrieval strategy to obtain candidate documents, generating a hybrid score in a weighted fusion manner, performing refined reordering on a preliminary candidate set by utilizing a cross encoder, then selecting high-confidence document fragments according to target culture, and performing high-confidence document segmentation on the selected high-confidence document fragments; and a cross-culture enhanced context prompt is constructed in combination with a culture constraint template and a soft prompt, so that a specific culture view angle and value preference can be automatically embedded in a large model in a generation process, and culture consistency contents are generated. The technical problem that a traditional large model is difficult to generate contents conforming to different cultural logics and expression habits under the same input is effectively solved, and the cultural logic accuracy and context adaptability of output contents are ensured.
Owner:LIAONING UNIVERSITY

Knowledge base Chinese document segmentation method and management system thereof

The invention discloses a knowledge base Chinese document segmentation method and a management system thereof, and the method comprises the following steps: S1, receiving a segmentation rule parameter set which is set by a user and comprises a segmentation separator, a segmentation maximum length value and a segmentation overlapping proportion; s2, analyzing the target document and extracting full-text content; s3, performing primary segmentation on the full-text content according to the segmentation separators to generate an initial segmentation set; and S4, performing iteration processing on each segment in the initial segment set. According to the method, a Chinese punctuation-driven integrity guarantee mechanism is created for the first time, full-angle full stop, question marks, exclamation marks, province marks and score marks are accurately positioned as sentence boundaries, and it is ensured that the head and the tail of each segment are complete Chinese sentences; designing an overlap transfer algorithm, extracting semantic segments containing Chinese punctuations from the tail of the segment, splicing the semantic segments to the head of the next segment, and constructing a context bridge between the segments; a pure Chinese punctuation adaptation system thoroughly eliminates segmentation errors caused by mixed use of Chinese and English.
Owner:GUANGDONG CHENGZHI TECH CO LTD