Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

157 results about "Text segmentation" patented technology

Text segmentation is the process of dividing written text into meaningful units, such as words, sentences, or topics. The term applies both to mental processes used by humans when reading text, and to artificial processes implemented in computers, which are the subject of natural language processing. The problem is non-trivial, because while some written languages have explicit word boundary markers, such as the word spaces of written English and the distinctive initial, medial and final letter shapes of Arabic, such signals are sometimes ambiguous and not present in all written languages.

Knowledge graph-based long text segmentation information retrieval result coherence enhancement method, system and equipment

The invention relates to the field of artificial intelligence, in particular to a long text segmentation information retrieval result coherence enhancement method, system and equipment based on a knowledge graph, and the method comprises the following steps: receiving a long text input by a user; segmenting the long text into a plurality of segments according to a preset length, and recording the original position and sequence of each segment; performing named entity recognition on each fragment, and extracting a key entity; querying related information in a knowledge graph according to the key entity, and obtaining attributes of the entity and a relationship between the attributes; integrating the related information obtained by query into the corresponding text fragment to form a context enhanced fragment; receiving a user query request; calculating a correlation score based on the context enhanced fragment and the user query, and retrieving a related text fragment by using a hybrid retrieval strategy; and selecting the related text fragment with the highest score to generate a final answer, and outputting the final answer to the user. Therefore, the information integrity and the context continuity in the retrieval process are ensured.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Text segmentation method and related equipment

PendingCN121328561AMathematical modelsSemantic analysisSemantic changeSemantic variation
The invention provides a text segmentation method and related equipment. The method comprises the steps of obtaining a to-be-processed text; segmenting the to-be-processed text into ordered statement sequences to obtain an initial statement set of the to-be-processed text; wherein the ordered statement sequence comprises a plurality of statements; the semantic variation, the confusion degree variation and the information entropy variation of a first target text block are calculated when a to-be-decided statement in an initial statement set of the to-be-processed text is added into the first target text block, and the first target text block is a set of multiple statements meeting a merging condition; determining the collaboration degree of the statement to be decided and the first target text block based on the semantic variable quantity, the confusion variable quantity and the information entropy variable quantity; and partitioning the to-be-processed text based on the collaboration degree of the to-be-decided statement and the first target text block to obtain a target partitioned text of the to-be-processed text. The text segmentation quality can be improved.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

AI classroom teaching quality evaluation method and system based on image recognition

The invention relates to the technical field, in particular to an AI classroom teaching quality evaluation method and system based on image recognition, and the method comprises the steps: obtaining a test paper RGB image and structured metadata, and generating multi-modal input data through graying, denoising and affine transformation; dynamically adjusting the weight through a double-branch feature extraction network (CNN + ViT), and extracting the global structure of the printed form and the local detail features of the handwritten form; generating a mixed font segmentation mask by using space attention, channel attention and dynamic routing regulation; optimizing handwritten answer recognition in combination with an educational knowledge graph; and fusing the answer expression and the attention data to generate a teaching quality evaluation report. The system comprises a data acquisition unit, a preprocessing unit, a feature extraction unit, a segmentation unit, an identification verification unit and an evaluation unit. According to the scheme, the accuracy of character segmentation and recognition in a mixed font scene is improved, and data support is provided for precise teaching in an education scene.
Owner:CHONGQING UNIV OF FINANCE & ECONOMICS

Long text news event element extraction method and system based on multi-agent collaboration

The invention relates to the technical field of information extraction, and discloses a long text news event element extraction method and system based on multi-agent collaboration, and the method comprises the following steps: constructing a large language model multi-agent system architecture; the central coordination agent receives the unstructured news text; the text segmentation agent outputs a logic unit and time description thereof; an event extraction agent extracts an event subject, an action and an object, and complements time and entity role information; the scene matching agent judges the fine-grained news scene category to which the scene matching agent belongs; the information extraction agent generates a structured event record; and the central coordination agent performs quality auditing and exception handling on each step of output to ensure the integrity and consistency of the output data. According to the method, the static relationship between the event main bodies can be constructed according to the structured event information through the relationship construction and context sorting agent, the event situation change is summarized based on the timeline, and the feasibility and the visualization of news event analysis of the system are improved.
Owner:CENT SOUTH UNIV

Engineering document generation method and system based on logic structure tree and LLM and medium

The invention discloses an engineering document generation method and system based on a logic structure tree and LLM and a medium. The method comprises the steps that original data are analyzed to form an initial text segment set X; calling LLM to generate an initial technical draft based on X, extracting a logic structure, and constructing an initial logic structure tree T; maintaining a search queue, for leaf nodes in the queue, acquiring context information related to semantics from X through a retrieval enhancement mechanism, after same-level brother node deduplication and global deduplication judgment, calling LLM to expand the nodes to generate child nodes, optimizing text description, and recursively enriching a tree structure; and finally, based on the X and the optimized tree, calling LLM to generate a title, an abstract and a background, generating content, details and key point chapters through traversal, and combining into a complete engineering document. The system comprises a document analysis module, a knowledge base module, a retrieval enhancement module, a logic structure tree maintenance module and a document generation module. The method is high in automation degree, and the generated document is clear in structure, complete in content and high in specialty.
Owner:SCHOOL OF SOFTWARE ZHEJIANG UNIV (NINGBO) MANAGEMENT CENT (NINGBO SOFTWARE EDUCATION CENT)

Archive keyword intelligent indexing method based on multi-modal semantic association

The invention discloses an intelligent archive keyword indexing method based on multi-modal semantic association, and relates to the technical field of archive keyword indexing. The method comprises the following steps: initially performing text segmentation of different target detection on text modal information; semantic expansion information block segmentation of corresponding target detection is carried out on image, audio and video modal information in an associated manner in sequence, and keyword extraction weight value statistics of positive samples with similar semantics and negative samples with opposite semantics is carried out on segmented sub-information blocks, and the segmented sub-information blocks are combined; the method comprises the following steps of: segmenting information blocks of image, audio and video modal information which are respectively used as initial processing objects in sequence, combining segmented text sub-information I, image sub-information II, audio sub-information III and video sub-information IV, and carrying out keyword extraction weight value statistics; and summarizing the test sample set obtained by statistics to realize keyword indexing through an established file keyword extraction model test. The keyword indexing accuracy can be improved.
Owner:贵阳市不动产登记中心

Contract review method based on large language model

PendingCN121095016AData processing applicationsSemantic analysisLinguistic modelArtificial intelligence and law
The invention relates to the cross technical field of artificial intelligence and law science and technology, and particularly discloses a contract review method based on a large language model, which comprises the following steps: obtaining a to-be-reviewed contract text, and carrying out analysis and formatting correction on the to-be-reviewed contract text to obtain the to-be-reviewed contract text after format correction; performing content correction on the to-be-examined contract text after the format correction so as to obtain a plurality of to-be-examined contract text segmentation blocks after the content correction; performing risk point review and clause level review on each to-be-reviewed contract text segmentation block after the content is corrected to obtain a risk point clause level review result corresponding to each to-be-reviewed contract text segmentation block after the content is corrected, and then generating a primary review risk point report of the to-be-reviewed contract text; and correcting the initial review risk point report of the contract text to be reviewed to output a final review risk point report of the contract text to be reviewed. According to the invention, the efficiency and accuracy of contract review can be improved.
Owner:无锡锡商银行股份有限公司

Text segmentation voice streaming processing method, system and device, medium and program product

The invention discloses a text segmentation voice streaming processing method, system and device, a medium and a program product, and the method comprises the steps: firstly, carrying out segmentation processing according to a semantic structure of a long text, and obtaining ordered text segments with complete semantics; speech synthesis is carried out on the first text segment to obtain a first audio segment, and SSE connection with the client is established while the first audio segment is synthesized; after the synthesis of the first audio clip is completed, the first audio clip is immediately transmitted to a segmentation buffer queue of the client through SSE connection, and meanwhile, speech synthesis is sequentially performed on subsequent text clips; and after the segmented buffering queue receives the first audio clip, the first audio clip is played immediately, and subsequent audio clips are continuously received and buffered while the first audio clip is played. According to the invention, the coordinated control of the AI speech synthesis progress and the streaming transmission progress is realized, the response speed of AI speech synthesis is obviously improved, and the switching fluency of segmented audios is effectively improved.
Owner:BEIJING DIANFU TECHNOLOGY CO LTD

Multi-body system dynamics intelligent modeling calculation system and method based on natural language

The invention discloses a multi-body system dynamics intelligent modeling calculation system and method based on a natural language, and belongs to the field of overall design of spacecrafts. The problems that in the prior art, when the simulation requirement related to extreme working conditions and strong nonlinear coupling is met, the period is long, manual experience is relied on, and the maintenance cost is high are solved. The system comprises a text segmentation embedding unit used for carrying out text segmentation embedding processing on a multi-body dynamics knowledge base and constructing a vectorized dynamics knowledge context; the large language model unit is used for receiving natural language task description input by a user, performing semantic understanding through LLM (Logistics Language Model) and generating a task analysis result according to dynamic knowledge context learning; and the task processing unit is used for classifying the tasks into inquiry tasks, command and control tasks or construction tasks according to the task analysis result, and calling corresponding modules according to LLM and MCP technologies. The method is used in the field of multi-body system dynamics modeling of spacecrafts.
Owner:HARBIN INST OF TECH

Knowledge base and knowledge graph dynamic access control method and system based on AI classification and grading

The invention discloses a knowledge base and knowledge graph dynamic access control method and system based on AI classification and grading, and relates to the technical field of data security and access control. Automatically classifying and grading text segments or knowledge graph entities by using an AI model through a data preprocessing module, generating structured tags, and binding the structured tags with original data; after the data are associated and stored to the database through the structured storage module, the authority management module maintains the accessible classification and grade range of the user; and the query filtering module calls the user permission and dynamically filters the retrieval result when the user retrieves, and only returns the authorized content. According to the method, automatic and fine-grained dynamic access control based on data content semantics is realized, the problems of dependence on manual labeling, coarse access control granularity and insufficient flexibility in the prior art are solved, data security and retrieval efficiency are considered, and the method is adaptive to multiple coincidence rules and standards.
Owner:DANGKANG DATA INTELLIGENCE TECHNOLOGY (GUANGDONG) CO LTD

Document retrieval and conflict detection system based on retrieval enhancement generation

The invention discloses a document retrieval and conflict detection system based on retrieval enhancement generation, and the system comprises a document loading module which is used for analyzing a multi-format document and extracting text content; the text segmentation module is used for dividing the text content into semantic segments; the vector index module is used for converting the semantic fragments into vectors and constructing indexes; the retrieval module is used for receiving user query and retrieving related semantic fragments; the conflict detection module is used for inputting the semantic fragments queried and retrieved by the user into a large language model for comparative analysis so as to identify logic or fact conflicts; the streaming output module is used for returning an intermediate reasoning process and a final conclusion of the large language model in real time; and the batch processing module is used for automatically processing the documents of the specified directory and batch query and generating a structured result file. According to the method, the system transparency and the user credibility are improved, and the robustness and efficiency in practical application are improved.
Owner:BEIJING TECH & BUSINESS UNIV

Streaming voice service concurrent processing method based on soft switch

The invention relates to a streaming voice service concurrent processing method based on soft switching, and the method comprises the steps: building and maintaining a global voice recognition channel with an ASR server, continuously forwarding a user voice stream to the ASR server during a whole session period, and receiving a recognition result; processing the recognition result by adopting a finite-state machine and a large language model to generate a response text for multiple rounds of dialogues; and concurrently executing TTS processing while the global speech recognition channel is continuously operated: segmenting a response text into a plurality of segments according to semantics to perform streaming speech synthesis, and immediately sending a pseudo response after receiving first speech data so as to trigger a soft switch system to pre-load a playing buffer area in advance. Through global reception and concurrent processing, smooth multi-round dialogues which can be interrupted at any time are realized, and the interaction delay is reduced from a second level to a millisecond level by using a text fragmentation and early media triggering technology, so that the user experience and the interaction efficiency of an intelligent call center are remarkably improved.
Owner:上海井星信息科技有限公司

Method and system for generating community policing risk compensation measures based on graph search enhancement generation

The application relates to the technical field of public security risks, and provides a community public security risk compensation measure generation method and system based on atlas retrieval enhancement generation. The method performs text segmentation and vectorization on obtained domain knowledge data in the field of public security risks, generates knowledge slices in the field of public security risks and corresponding knowledge vectors, and stores the knowledge slices and the knowledge vectors into a vector database; meanwhile, a domain knowledge atlas in the field of community public security risks is constructed, and an index relationship between the domain knowledge atlas and the knowledge vectors is stored in the vector database, so as to construct a domain knowledge base in the field of public security risks; then, a user risk query is analyzed through a large language model, risk elements related to the user risk query are searched in the domain knowledge atlas, and knowledge slices corresponding to the risk elements are searched in the domain knowledge base according to the index relationship, so as to generate public security risk compensation measures corresponding to the user risk query.
Owner:CITIC GUOAN INFORMATION TECH +1

Multi-agent cooperation-based long text news event element extraction method and system

This invention relates to the field of information extraction technology, and discloses a method and system for extracting elements from long-text news events based on multi-agent collaboration. The method includes the following steps: constructing a large language model multi-agent system architecture; a central coordinating agent receiving unstructured news text; a text segmentation agent outputting logical units and their time descriptions; an event extraction agent extracting the event subject, action, and object, and completing time and entity role information; a scene matching agent determining the fine-grained news scene category to which it belongs; an information extraction agent generating structured event records; and a central coordinating agent performing quality review and anomaly handling on each step of the output to ensure the integrity and consistency of the output data. This invention, through relationship building and contextual analysis agents, can construct static relationships between event subjects based on structured event information, and summarize event situation changes based on a timeline, improving the feasibility and visualization of the system's news event analysis.
Owner:CENT SOUTH UNIV

Commodity label generation method and device for e-commerce scene, and storage medium

The invention relates to an e-commerce scene-oriented commodity label generation method and device, and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the steps of obtaining native text information and an image set of a target commodity; performing information identification and extraction on the image set, converting related visual data into structured text description data, and fusing the structured text description data into a to-be-processed text corpus with multi-dimensional information; dynamically generating a theme label structure for the target commodity; performing text segmentation on the to-be-processed text corpus according to the topic tag structure to obtain grouped texts corresponding to different topic tags, and feeding back each grouped text to a corresponding topic expert sub-model to perform candidate tag extraction; and semantic fusion is carried out, and a fused label set is used as a final label of the target commodity. According to the method, the integrity, accuracy and robustness of commodity label generation in an e-commerce scene are remarkably improved, and the limitation of dependence on manpower or a fixed label library is reduced.
Owner:HANGZHOU NO TABLE ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

A text segmentation and sensitive word detection method based on matrix multiplication

ActiveCN115757721BEnergy efficient computingText database indexingAlgorithmDeterministic finite automaton
A text segmentation and sensitive word detection method based on matrix multiplication, comprising the following steps: obtaining an original text string and a sensitive word library; constructing a deterministic finite state automaton tree diagram of sensitive words according to the sensitive word library; converting the original text string into a text character two-dimensional matrix and recording the length of the text character two-dimensional matrix; constructing a matching two-dimensional matrix according to a horizontal matching rule, a vertical matching rule, an oblique matching rule and an inverse oblique matching rule, the length of the matching two-dimensional matrix being the same as that of the text character two-dimensional matrix; performing dot multiplication processing on the text character two-dimensional matrix and the matching two-dimensional matrix to obtain a corresponding result matrix; generating a corresponding matching text string according to the result matrix, and matching with the deterministic finite state automaton tree diagram to determine whether there is a sensitive word. The application supports interval text character information detection, improves the sensitive word detection accuracy and detection efficiency.
Owner:XUANCAI INTERACTIVE NETWORK SCI & TECH

A text segmentation method, system, computer device, and storage medium

This invention relates to the field of artificial intelligence technology, providing a text segmentation method, system, computer device, and storage medium, comprising: acquiring the text to be segmented; dividing each sentence of the text into multiple semantic blocks according to a left-to-right order, with the end point as the boundary; if two consecutive characters do not have a connected record in the vocabulary of the text to be segmented, then the preceding character of the two consecutive characters is recorded as the end point; performing full segmentation on each semantic block to obtain all possible segmentation methods for that semantic block; calculating the probability of traversing all segmentation methods for each semantic block in a left-to-right order, and selecting the segmentation method with the highest probability as the final segmentation result. The segmentation scheme of this invention is based on probability, traverses all solutions within a block, and comprehensively considers the context of the text, resulting in more accurate segmentation results, reducing labor costs, and improving segmentation accuracy.
Owner:ONE CONNECT SMART TECH CO LTD SHENZHEN

LLM-based power field intelligent question answering method, system, device and medium

The application discloses a power field intelligent question and answer method, system, device and medium based on LLM, wherein the method comprises the following steps: S1: obtaining text content by reading various documents in the power field, obtaining text content blocks through text segmentation, constructing and saving context association information of the documents, vectorizing the text content blocks by using a vector model, and saving the text content blocks and word vectors to a vector database; S2: performing multi-path recall and retrieval summarization according to a user question, and rearranging the retrieved documents; S3: finding context document records of a top1 document obtained through rearrangement in the multi-path recall stage by using a context retriever, and merging the context document records into a complete document; and S4: combining the complete document content and the user question to input a prompt into an LLM model to answer the question. The application overcomes the limitation of a semantic-based retrieval mode, and provides a more accurate, comprehensive and adaptable intelligent question and answer solution.
Owner:FUJIAN YIRONG INFORMATION TECH +1

Big language model data illusion identification method and device and storage medium

The invention provides an illusion recognition method and device for big language model data and a storage medium, and belongs to the technical field of computers, and the method comprises the steps: generating a plurality of first verification questions based on segmented texts corresponding to all knowledge themes in a knowledge base of a target big model; generating a plurality of answers for each first verification question; based on the text segment and the plurality of answers, calling a preset contradiction evaluation large model to perform consistency evaluation on each answer of the first verification question to obtain a respective corresponding first evaluation result; based on each first evaluation result, determining a target verification question from the plurality of first verification questions, and generating a second verification question for an answer corresponding to the target verification question; performing contradictory evaluation based on the second verification question and the answer to obtain respective corresponding second evaluation results; and determining the illusion recognition result of the knowledge base based on the first evaluation result or the second evaluation result. The illusion recognition accuracy and recognition efficiency of the large model data are improved.
Owner:WUHAN BIG PULP IND DEV CO LTD

Data retrieval method, server, terminal, and storage medium

The application discloses a data retrieval method, a server, a terminal and a storage medium, which are used for providing a knowledge base question and answer retrieval function for VDI users through a server without occupying the storage space of the server, and providing a safer, more reliable and more intelligent question and answer retrieval scheme. The method comprises the following steps: acquiring first text information from a first remote terminal in one or more remote terminals, wherein the first text information corresponds to a first embedding vector; if the similarity of the first embedding vector and a second embedding vector in a vector database is greater than or equal to a threshold value, determining a retrieval result of the first text information according to the second text information corresponding to the second embedding vector and the first embedding vector, wherein the second text information is obtained by text segmentation according to source documents of the one or more remote terminals.
Owner:RUIJIE NETWORKS CO LTD

Infinite context large model processing method and device based on lexical element memory

The invention discloses an infinite context large model processing method and device based on lexical element memory. The method comprises the following steps: firstly, segmenting an ultra-long text into text paragraphs with fixed lengths; secondly, constructing a local attention module based on a decoder type Transform, calculating text paragraphs, and outputting a local attention result and a related vector; and thirdly, constructing a token lexical element memory pool to maintain each lexical element memory unit, and calculating cross-paragraph global attention and outputting a result in combination with the query vector. And fourthly, fusing local and global attention results to realize an infinite context. And 5, based on the correlation vector, incrementally updating the lexical token memory pool and the normalization item, and providing support for subsequent paragraph processing. And 6, constructing a paragraph memory module, and storing the hidden state of the last lexical element of the paragraph for generating the first lexical element of the next text paragraph. The problems that in the prior art, when super-long texts are processed, calculation and storage expenses are large, and semantic information is prone to being lost are solved.
Owner:烟台国工智能科技有限公司

A method, system, and device for text segmentation and highlighting based on speech character position mapping

PendingCN122309020APattern recognitionCharacter (computing)
This invention discloses a method, system, and device for text segmentation highlighting based on speech character position mapping, applied in the field of data processing technology. It achieves text segmentation highlighting through scrolling character position mapping. First, the text content, scrolling sequence, speed, and font layout are preprocessed and uniformly converted into character coordinates and progress percentage format. Character indices and highlight intervals are extracted through position mapping and segmentation detection to form display baseline information. Based on the scrolling-driven calculation module, text features are modulated and non-viewport content is filtered. A highlighting following effect is generated after pixel restoration and synchronization calibration. The highlighting module is directly connected to a fixed-logic prompting engine, and synchronization calibration of scrolling and chapter modes is only performed at the front-end mapping layer, achieving parameter optimization. Seamless adaptation of the rendering framework is achieved through dimensional alignment. Using the prompting engine as a timing discriminator, a complete technical solution for low-computing-power, high-precision, and pluggable scrolling text synchronization and segmentation highlighting is ultimately formed.
Owner:ZHANGZHOU SEETEC OPTOELECTRONICS TECH CO LTD

A text image detection method, device, medium and equipment

Embodiments of the present application disclose a text image detection method, device, medium and equipment. The method comprises: obtaining a text image to be detected; using a pre-trained text detection model to detect the text image to be detected to obtain a text detection result; the text detection model comprises a backbone network, a region detection branch and a text segmentation branch; the region detection branch and the text segmentation branch are both connected after the backbone network. The technical solution can enhance the generalization ability of the text detection model by building a multi-branch text detection model, realize accurate region detection frame regression and text segmentation, and effectively improve the robustness and accuracy of text image detection and reduce the detection time.
Owner:CHINA POST INFORMATION TECH (BEIJING CO LTD

A method and system for constructing a large language model enhanced emission trading knowledge graph

The application discloses a kind of big language model enhanced emission trading knowledge graph construction method and system, it is related to machine learning field.The present application obtains the unstructured text data in the field of emission trading, carries out document analysis and text segmentation to text data, obtains standardized text fragment;Based on the field of emission trading adaptability prompt engineering guides big language model, carries out entity recognition and relationship extraction to the text fragment, obtains initial knowledge triple and processing, obtains high-quality knowledge triple;Text fragment is converted and constructs structured triple with high-quality knowledge triple and semantic vector fusion's bimodal knowledge representation;The bimodal knowledge representation is stored to graph database, and constructs the field of emission trading knowledge graph;Based on double-path retrieval, carry out retrieval and knowledge traceability to knowledge graph, realize the field of emission trading intelligent question and answer and knowledge application.The present application realizes the efficient conversion from unstructured data to structured knowledge.
Owner:CHINA JAPAN FRIENDSHIP ENVIRONMENTAL PROTECTION CENT

Translation method and translation system

The invention provides a translation method and a translation system. According to the translation method, firstly, a predefined text knowledge base is adopted to preprocess text segments needing to be precisely processed in a to-be-translated text, so that a prompt of an AI large language model contains a consistent text comparison table customized for the to-be-translated text; wherein fixed translation of specific text segments (such as proper nouns, phrases or sentences) in a to-be-translated text in a target language is defined; and translating the preprocessed text by adopting the AI large language model. The long text such as the novel can be segmented into chapters and sections by adopting an automatic text segmentation technology before translation, and the translated texts of all the chapters and sections are recombined into a target file. According to the technical scheme, the automatic text segmentation technology and a cooperation mechanism of text consistency preprocessing and comparison table guiding type Prompt guidance are combined, the problems of key expression inconsistency, semantic drift and the like occurring in long text translation are effectively avoided, and the continuity and readability of translated texts are improved.
Owner:SHANGHAI CANFENG NETWORK TECHNOLOGY CO LTD

Network security operation aid decision-making method, system and device based on AIGC and BERT

The invention discloses a network security operation auxiliary decision-making method, system and device based on AIGC and BERT, and relates to the technical field of network security, and the method comprises the steps: extracting an attribute field of a security operation case library; performing text segmentation on the enriched safety operation case information; through a preprocessing process, mounting of a safety operation case library is completed, decision generation and evaluation are carried out, and a decision case is generated. According to the method, the problems that AI-driven attacks are difficult to identify and response lags are solved efficiently, the alarm research and judgment accuracy and efficiency are improved, the expert resource pressure is relieved, and reliable support is provided for quick response and full-automatic disposal of high-adversarial attacks.
Owner:CHINA ZHESHANG BANK +1

Long text adaptive segmentation method suitable for information extraction

The invention relates to the field of natural language processing, and provides a long text adaptive segmentation method suitable for information extraction in order to realize text segmentation based on semantic equilibrium, which comprises the following steps of: uniformly slicing a long text, calculating a semantic embedding vector of each slice, and simultaneously calculating semantic embedding vectors of a plurality of target fields; calculating semantic similarity between each target field and each slice, and normalizing the similarity into discrete probability distribution; constructing a field distribution difference matrix based on divergence between semantic density distributions of the fields, and aggregating the fields with similar semantic density distributions into a field cluster by adopting a connected graph clustering mode based on threshold control; and carrying out continuous aggregation on the initial slices according to a preset target segmentation number on the basis of the average semantic similarity of the field clusters on the slices. According to the method, the dynamic segmentation of the long text can be realized according to the field semantic features, and the positioning precision and the processing efficiency of the information extraction task are improved.
Owner:CHINA HYDROELECTRIC ENGINEERING CONSULTING GROUP CHENGDU RESEARCH HYDROELECTRIC INVESTIGATION DESIGN AND INSTITUTE

Aero-engine test document automatic generation method and system based on demand analysis

The invention belongs to the technical field of artificial intelligence and aviation engineering crossing, and provides an aero-engine test document automatic generation method and system based on demand analysis. Segmenting the test technical requirement document by adopting a hybrid text segmentation algorithm based on semantic vector similarity, and constructing a bimodal storage structure; high-precision fusion of multiple paths of retrieval results is realized by utilizing a reciprocal ranking fusion algorithm and a weighted reciprocal ranking fusion algorithm, and a high-correlation context is extracted; structured test points are automatically analyzed and generated based on a fine-tuning large model, and multi-target task grouping is carried out in combination with test rules such as a safety envelope, positive and negative temperature isolation and single-day duration; and checking and outputting compliance test task groups through a rule-driven post-processing mechanism. According to the method, the production scheduling efficiency and the resource utilization rate are remarkably improved, dynamic adjustment is supported, and the method is suitable for multi-equipment and high-complexity aero-engine test scenes.
Owner:AECC SICHUAN GAS TURBINE RES INST

A fault library modeling method, device and product based on a decision chain graph

The application discloses a fault library modeling method and device based on a decision chain graph, and a product; the method comprises the following steps: obtaining structured account data and unstructured text data; performing data format cleaning and text paragraph segmentation on the unstructured text data to obtain text segmentation segments; performing sentence-level standardization processing on natural language fields in the text segmentation segments and the structured account data to obtain target sentence-level text; identifying key semantic unit information from the target sentence-level text, the key semantic unit information being used for describing a decision unit four-tuple; performing directed chain combination, chain node merging and standardization processing on the decision unit four-tuple to obtain a plurality of decision chains, and generating a decision chain graph based on the plurality of decision chains, so as to model a subway power supply fault through the decision chain graph. The embodiment of the application can make knowledge have a decision chain structure of'state-judgment-action-result', thereby realizing flow modeling and intelligent response of the fault library.
Owner:BEIJING MASS TRANSIT RAILWAY OPERATION CORPORATION LIMITED

Data depolarization and alignment enhancement method for large language model, electronic equipment and storage medium

The invention belongs to the technical field of generative artificial intelligence, and provides a data depolarization and alignment enhancement method for a large language model, electronic equipment and a storage medium. The method comprises the steps of general text collection, single sentence disassembly, template matching, emotion labeling, object conflict group division, prejudice text and neutral text division, neutral dialogue model construction, prejudice model construction, synchronous answer generation, double-model output difference degree calculation, answer examination and neutral answer output. According to the method, through template matching, emotion labeling and conflict group division, complex cross bias detection is converted into objects, and emotion comparison statistical analysis is performed, so that the depolarization complexity is reduced, and the recognition efficiency of bias texts is improved; by constructing the neutral dialogue model and the prejudice model, error reference is provided, the process that traditional human preference alignment needs to collect a large number of personnel evaluations is avoided, and the training of the neutral dialogue model and the examination process during work are optimized.
Owner:GUANGXI ACAD OF SCI