Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

238 results about "Text segmentation" patented technology

Text segmentation is the process of dividing written text into meaningful units, such as words, sentences, or topics. The term applies both to mental processes used by humans when reading text, and to artificial processes implemented in computers, which are the subject of natural language processing. The problem is non-trivial, because while some written languages have explicit word boundary markers, such as the word spaces of written English and the distinctive initial, medial and final letter shapes of Arabic, such signals are sometimes ambiguous and not present in all written languages.

Basic-level power supply enterprise compliance risk early warning system and method based on big data analysis

The invention discloses a grassroots power supply enterprise compliance risk intelligent system and method based on big data analysis. The data acquisition unit is used for acquiring business operation data, historical violation records and policy and regulation update data of basic power supply enterprises to form a unified compliance data set. And the natural language processing unit performs text word segmentation and correlation analysis on the policy and regulation and violation record data, extracts key risk factors and labels compliance risk labels. And the risk feature construction unit performs multi-dimensional feature fusion on the business operation data and the compliance risk label data to generate a feature matrix for risk identification. And the intelligent risk assessment unit performs real-time analysis on the feature matrix by using a pre-trained machine learning model, identifies compliance risk categories and levels, and generates early warning information. According to the method, the accuracy of compliance risk identification is improved by using big data analysis and an intelligent algorithm, the compliance management cost of basic-level power supply enterprises is reduced, and the operation safety and compliance of the enterprises are improved.
Owner:JURONG CITY POWER SUPPLY BRANCH OF STATE GRID JIANGSU ELECTRIC POWER CO LTD +1

Semantic enhancement adaptive partitioning method and system for natural resource large model questions and answers

The invention provides a semantic enhancement adaptive partitioning method and system for natural resource large model questions and answers, and aims to solve the problems of difficulty in term boundary recognition, damage to semantic integrity and the like. According to the method, three core technologies including theme perception coarse-grained paragraph division, self-adaptive sliding window theme hierarchy division and embedded perception context self-adaptive text segmentation are fused. The system firstly analyzes a natural resource long text structure, identifies titles and theme levels and aligns associated contents; paragraphs are extracted according to a theme perception strategy and are subdivided into sentence sets according to grammar rules; an improved sliding window mechanism is adopted to divide sentences into window sentence block groups. The method is characterized in that a dynamic aggregation threshold mechanism is introduced, the semantic association degree between adjacent sentence blocks is calculated through an embedded perception context semantic segmentation technology, whether the sentence blocks are combined or not is judged by combining a similarity distribution change trend and a dynamic adjustment threshold, self-adaptive delimitation of semantic boundaries is achieved, and text blocks which are clear in structure and coherent in semantics are generated.
Owner:HUBEI PROVINCIAL DEPT OF NATURAL RESOURCES INFORMATION CENT +1

Heterogeneous data-driven marketing channel clustering modeling and optimizing method

The invention discloses a marketing channel clustering modeling and optimizing method driven by heterogeneous data. The method comprises the steps of obtaining a multi-dimensional data set of a marketing channel, and performing preprocessing; the data set comprises structured data and unstructured data; and extracting a text word segmentation result and an emotion feature from the unstructured data, and extracting an image feature to generate a structured feature set, an unstructured feature set and the like. According to the method, structured and unstructured data can be integrated, and potential information of marketing channels can be comprehensively mined. Through a deep learning model (an auto-encoder, a variational auto-encoder and a generative adversarial network), features are extracted and optimized layer by layer, and the quality of the features and the performance of the model are improved. And in combination with a clustering algorithm and an evaluation index, the accuracy and reliability of a clustering result are ensured, and powerful support is provided for optimization of a marketing channel.
Owner:FUJIAN CHUANZHENG COMM COLLEGE

Professional knowledge graph construction method for medical field

The invention relates to a professional knowledge graph construction method for the medical field, and the method comprises the following steps: obtaining a multi-modal medical document, carrying out the deep analysis through a Transform-based multi-modal model, and obtaining an analyzed multi-modal medical document; dynamically segmenting the analyzed multi-modal medical document by adopting a context-based adaptive text segmentation method to obtain independent text blocks; based on the independent text blocks, performing structured extraction, and performing verification by adopting a dynamic dual-verification mechanism to form knowledge sub-graphs; and based on the knowledge sub-graph, further constructing a standardized medical knowledge graph. Compared with the prior art, the method has the advantages of high precision, expandability and the like.
Owner:SHANGHAI JIAOTONG UNIV

Knowledge graph-based long text segmentation information retrieval result coherence enhancement method, system and equipment

The invention relates to the field of artificial intelligence, in particular to a long text segmentation information retrieval result coherence enhancement method, system and equipment based on a knowledge graph, and the method comprises the following steps: receiving a long text input by a user; segmenting the long text into a plurality of segments according to a preset length, and recording the original position and sequence of each segment; performing named entity recognition on each fragment, and extracting a key entity; querying related information in a knowledge graph according to the key entity, and obtaining attributes of the entity and a relationship between the attributes; integrating the related information obtained by query into the corresponding text fragment to form a context enhanced fragment; receiving a user query request; calculating a correlation score based on the context enhanced fragment and the user query, and retrieving a related text fragment by using a hybrid retrieval strategy; and selecting the related text fragment with the highest score to generate a final answer, and outputting the final answer to the user. Therefore, the information integrity and the context continuity in the retrieval process are ensured.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Text segmentation method and related equipment

PendingCN121328561AMathematical modelsSemantic analysisSemantic changeSemantic variation
The invention provides a text segmentation method and related equipment. The method comprises the steps of obtaining a to-be-processed text; segmenting the to-be-processed text into ordered statement sequences to obtain an initial statement set of the to-be-processed text; wherein the ordered statement sequence comprises a plurality of statements; the semantic variation, the confusion degree variation and the information entropy variation of a first target text block are calculated when a to-be-decided statement in an initial statement set of the to-be-processed text is added into the first target text block, and the first target text block is a set of multiple statements meeting a merging condition; determining the collaboration degree of the statement to be decided and the first target text block based on the semantic variable quantity, the confusion variable quantity and the information entropy variable quantity; and partitioning the to-be-processed text based on the collaboration degree of the to-be-decided statement and the first target text block to obtain a target partitioned text of the to-be-processed text. The text segmentation quality can be improved.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

AI classroom teaching quality evaluation method and system based on image recognition

The invention relates to the technical field, in particular to an AI classroom teaching quality evaluation method and system based on image recognition, and the method comprises the steps: obtaining a test paper RGB image and structured metadata, and generating multi-modal input data through graying, denoising and affine transformation; dynamically adjusting the weight through a double-branch feature extraction network (CNN + ViT), and extracting the global structure of the printed form and the local detail features of the handwritten form; generating a mixed font segmentation mask by using space attention, channel attention and dynamic routing regulation; optimizing handwritten answer recognition in combination with an educational knowledge graph; and fusing the answer expression and the attention data to generate a teaching quality evaluation report. The system comprises a data acquisition unit, a preprocessing unit, a feature extraction unit, a segmentation unit, an identification verification unit and an evaluation unit. According to the scheme, the accuracy of character segmentation and recognition in a mixed font scene is improved, and data support is provided for precise teaching in an education scene.
Owner:CHONGQING UNIV OF FINANCE & ECONOMICS

Text segmentation method, question answering method and related devices and equipment

The invention discloses a text segmentation method, a question answering method and related devices and equipments.The text segmentation method comprises the steps that the confusion degree of each sentence text in a to-be-processed text is obtained; for each sentence text, obtaining a confusion degree comparison result between the sentence text and the adjacent sentence text, and determining whether to select the sentence text as a segmentation point of the to-be-processed chapter based on the comparison result; and segmenting the to-be-processed chapter based on each segmentation point to obtain a plurality of text blocks. According to the scheme, the segmentation adaptivity of the chapter text can be improved, and the semantic and logic coherence of the text blocks obtained through segmentation can be improved.
Owner:HEFEI IFLY DIGITAL TECH CO LTD

Long text news event element extraction method and system based on multi-agent collaboration

The invention relates to the technical field of information extraction, and discloses a long text news event element extraction method and system based on multi-agent collaboration, and the method comprises the following steps: constructing a large language model multi-agent system architecture; the central coordination agent receives the unstructured news text; the text segmentation agent outputs a logic unit and time description thereof; an event extraction agent extracts an event subject, an action and an object, and complements time and entity role information; the scene matching agent judges the fine-grained news scene category to which the scene matching agent belongs; the information extraction agent generates a structured event record; and the central coordination agent performs quality auditing and exception handling on each step of output to ensure the integrity and consistency of the output data. According to the method, the static relationship between the event main bodies can be constructed according to the structured event information through the relationship construction and context sorting agent, the event situation change is summarized based on the timeline, and the feasibility and the visualization of news event analysis of the system are improved.
Owner:CENT SOUTH UNIV

Supply chain fraud behavior early warning method and device based on large language model

The invention discloses a supply chain fraud behavior early warning method and device based on a large language model, and relates to the field of data analysis, and the method comprises the steps: obtaining supply chain multi-source data, and carrying out the processing of the supply chain multi-source data according to the data type; in the fine tuning process, a LoRA module is injected into a linear layer of a pre-trained large language model base, and the rank value of the LoRA module is dynamically adjusted according to the gradient; according to the text word segmentation data, the entity type, the aligned historical order data, the aligned historical logistics data and the dynamic space-time diagram, constructing a multi-modal prefix guide vector; the multi-modal prefix guide vector and an original input sequence of a linear layer of a pre-trained large language model base are spliced and then input into the linear layer of the pre-trained large language model base, a supply chain fraud behavior early warning model subjected to fine adjustment is obtained through fine adjustment, and early warning is conducted on supply chain fraud behaviors of related suppliers. According to the method, the problems of low supply chain fraud behavior identification accuracy, large training parameters and the like in the prior art are solved.
Owner:XIAMEN MEIYA YIAN INFORMATION TECH CO LTD

Text block semantic coherence detection method oriented to RAG question and answer system

The invention belongs to the field of artificial intelligence technology machine learning and semantic understanding, and relates to semantic analysis and semantic recognition, in particular to a text block semantic coherence detection method oriented to an RAG question and answer system, which specifically quantifies semantic loss of different blocks, assists in rejecting sampling to construct a supervised fine tuning (SFT) and reinforcement learning (RL) training set, and improves the semantic coherence of the different blocks. Therefore, the text partitioning capability of the large language model is improved, and the retrieval effect and the answer generation quality of the RAG system are optimized. According to the method, the semantic consistency in the blocks and the semantic jump degree at the boundaries of the blocks are quantified, an objective basis is provided for evaluating the quality of the text blocks, the effectiveness of the algorithm is verified through design experiments, and the method has important significance in improving the text blocking capacity of a large language model, the retrieval effect of an RAG question and answer system and the answer generation quality.
Owner:GUIZHOU NORMAL UNIVERSITY

Cultural resource digitization implementation method based on image recognition

The embodiment of the invention provides a cultural resource digitization implementation method based on image recognition, and the method comprises the steps: extracting a text candidate region in an enhanced image based on text edge positioning and non-text region filtering, carrying out the text segmentation of the text candidate region through employing a DB segmentation algorithm, carrying out the segmentation processing of a text line image obtained through segmentation, and carrying out the segmentation processing of the text line image. A coincident text segment exists between the segment tail of the first text segment and the segment head of the second text segment; performing text recognition on the first text segment and the second text segment by using a text recognition model, and using an obtained recognition result of the coincident text segment for optimization adjustment of the text recognition model; and taking a new to-be-recognized image as input of the optimized and adjusted text recognition model to obtain a text recognition result of the current to-be-recognized image, and constructing and outputting a culture resource database containing the text recognition result. According to the embodiment of the invention, image information such as ancient books and literatures can be digitalized, and a culture resource database convenient to manage and apply is constructed.
Owner:时代新媒体出版社有限责任公司

Method and system for generating hospitalization first-time disease course record based on large model

The invention provides a hospitalization first-time disease course record generation method and system based on a large model, and the method comprises the steps: carrying out the text segmentation, format conversion, term standardization and theme recognition through employing a historical hospitalization document record, obtaining semantic transcription data through the similarity, building a mapping relation, and carrying out the calculation of the semantic transcription data. According to the method, the medical record documents called and stored in the electronic medical record system are processed, semantic transcription data corresponding to the medical record documents are obtained, and case medical records with similarity meeting a threshold value and stored in the electronic medical record system are matched through keyword and semantic mixed retrieval according to the medical record documents and the semantic transcription data. According to the semantic transcription data and the first course record initial cue word, constructing an optimized cue word with department and disease type characteristics, and according to the optimized cue word, the semantic transcription data and the case medical record, inputting the optimized cue word, the semantic transcription data and the case medical record into a hospitalization first course record generation model trained by adopting historical data to finally obtain a hospitalization first course record. The problem that in the prior art, the accuracy of the first record content is low is solved.
Owner:ZHONGSHAN HOSPITAL FUDAN UNIV

Engineering document generation method and system based on logic structure tree and LLM and medium

The invention discloses an engineering document generation method and system based on a logic structure tree and LLM and a medium. The method comprises the steps that original data are analyzed to form an initial text segment set X; calling LLM to generate an initial technical draft based on X, extracting a logic structure, and constructing an initial logic structure tree T; maintaining a search queue, for leaf nodes in the queue, acquiring context information related to semantics from X through a retrieval enhancement mechanism, after same-level brother node deduplication and global deduplication judgment, calling LLM to expand the nodes to generate child nodes, optimizing text description, and recursively enriching a tree structure; and finally, based on the X and the optimized tree, calling LLM to generate a title, an abstract and a background, generating content, details and key point chapters through traversal, and combining into a complete engineering document. The system comprises a document analysis module, a knowledge base module, a retrieval enhancement module, a logic structure tree maintenance module and a document generation module. The method is high in automation degree, and the generated document is clear in structure, complete in content and high in specialty.
Owner:SCHOOL OF SOFTWARE ZHEJIANG UNIV (NINGBO) MANAGEMENT CENT (NINGBO SOFTWARE EDUCATION CENT)

Orthopedic nursing intelligent data processing system based on cloud platform

The invention provides an orthopedic nursing intelligent data processing system based on a cloud platform, and relates to the field of intelligent data processing.The orthopedic nursing intelligent data processing system comprises the steps that firstly, a newly-added nursing document is transmitted to a temporary storage area of a cloud end and read, and then format cleaning and text segmentation are conducted on the nursing document to generate sequence distribution described by sentence granularity; then, the sequence is distributed and input into an NLP model deployed at a cloud end to complete named entity recognition, relation extraction, negative word detection and attribute extraction to obtain structured information of the nursing document, and then the structured information is subjected to standardization and duplicate removal processing to generate standard structured information of the nursing document; and the information is stored in a structured database of the cloud. In this way, the efficiency and accuracy of nursing data processing can be improved, and the deep value of data can be released.
Owner:SHANGHAI PUDONG HOSPITAL

Archive keyword intelligent indexing method based on multi-modal semantic association

The invention discloses an intelligent archive keyword indexing method based on multi-modal semantic association, and relates to the technical field of archive keyword indexing. The method comprises the following steps: initially performing text segmentation of different target detection on text modal information; semantic expansion information block segmentation of corresponding target detection is carried out on image, audio and video modal information in an associated manner in sequence, and keyword extraction weight value statistics of positive samples with similar semantics and negative samples with opposite semantics is carried out on segmented sub-information blocks, and the segmented sub-information blocks are combined; the method comprises the following steps of: segmenting information blocks of image, audio and video modal information which are respectively used as initial processing objects in sequence, combining segmented text sub-information I, image sub-information II, audio sub-information III and video sub-information IV, and carrying out keyword extraction weight value statistics; and summarizing the test sample set obtained by statistics to realize keyword indexing through an established file keyword extraction model test. The keyword indexing accuracy can be improved.
Owner:贵阳市不动产登记中心

Contract review method based on large language model

PendingCN121095016AData processing applicationsSemantic analysisLinguistic modelArtificial intelligence and law
The invention relates to the cross technical field of artificial intelligence and law science and technology, and particularly discloses a contract review method based on a large language model, which comprises the following steps: obtaining a to-be-reviewed contract text, and carrying out analysis and formatting correction on the to-be-reviewed contract text to obtain the to-be-reviewed contract text after format correction; performing content correction on the to-be-examined contract text after the format correction so as to obtain a plurality of to-be-examined contract text segmentation blocks after the content correction; performing risk point review and clause level review on each to-be-reviewed contract text segmentation block after the content is corrected to obtain a risk point clause level review result corresponding to each to-be-reviewed contract text segmentation block after the content is corrected, and then generating a primary review risk point report of the to-be-reviewed contract text; and correcting the initial review risk point report of the contract text to be reviewed to output a final review risk point report of the contract text to be reviewed. According to the invention, the efficiency and accuracy of contract review can be improved.
Owner:无锡锡商银行股份有限公司

Text segmentation voice streaming processing method, system and device, medium and program product

The invention discloses a text segmentation voice streaming processing method, system and device, a medium and a program product, and the method comprises the steps: firstly, carrying out segmentation processing according to a semantic structure of a long text, and obtaining ordered text segments with complete semantics; speech synthesis is carried out on the first text segment to obtain a first audio segment, and SSE connection with the client is established while the first audio segment is synthesized; after the synthesis of the first audio clip is completed, the first audio clip is immediately transmitted to a segmentation buffer queue of the client through SSE connection, and meanwhile, speech synthesis is sequentially performed on subsequent text clips; and after the segmented buffering queue receives the first audio clip, the first audio clip is played immediately, and subsequent audio clips are continuously received and buffered while the first audio clip is played. According to the invention, the coordinated control of the AI speech synthesis progress and the streaming transmission progress is realized, the response speed of AI speech synthesis is obviously improved, and the switching fluency of segmented audios is effectively improved.
Owner:BEIJING DIANFU TECHNOLOGY CO LTD

Multi-body system dynamics intelligent modeling calculation system and method based on natural language

The invention discloses a multi-body system dynamics intelligent modeling calculation system and method based on a natural language, and belongs to the field of overall design of spacecrafts. The problems that in the prior art, when the simulation requirement related to extreme working conditions and strong nonlinear coupling is met, the period is long, manual experience is relied on, and the maintenance cost is high are solved. The system comprises a text segmentation embedding unit used for carrying out text segmentation embedding processing on a multi-body dynamics knowledge base and constructing a vectorized dynamics knowledge context; the large language model unit is used for receiving natural language task description input by a user, performing semantic understanding through LLM (Logistics Language Model) and generating a task analysis result according to dynamic knowledge context learning; and the task processing unit is used for classifying the tasks into inquiry tasks, command and control tasks or construction tasks according to the task analysis result, and calling corresponding modules according to LLM and MCP technologies. The method is used in the field of multi-body system dynamics modeling of spacecrafts.
Owner:HARBIN INST OF TECH

Traffic transportation government affair hotline mining method based on natural language processing

The invention discloses a traffic transportation government affair hotline mining method based on natural language processing, and belongs to the technical field of artificial intelligence and machine learning. According to the method, for structured and unstructured data in a traffic transportation government affair hotline, firstly, a five-class word segmentation dictionary containing cleaning words, noise words, synonyms, additive words and stop words is constructed; converting the unstructured text into structured data through improved data cleaning, text word segmentation and feature representation; carrying out clustering analysis by utilizing an LDA topic model, and extracting public demand topics and high-frequency keywords; hot appeals and trend changes are mined in combination with space-time analysis and association analysis, and finally a visual analysis report is generated. According to the method, the problems of poor model interpretability and insufficient field adaptability in the prior art are solved, the accuracy of hotline data processing and the effectiveness of theme recognition are remarkably improved through multi-dictionary collaborative optimization and field knowledge fusion, and accurate decision support is provided for a traffic transportation management department. A real taxi field case in a certain city is used as an example for research, an experiment proves that the method has an accurate theme identification function, and the complaint and report work order amount in the traffic transportation government affair hotline taxi field in the city is reduced by 20% on year-on-year basis in 2024.
Owner:乌若愚

Method and system for improving text content recognition and classification and computer equipment

The invention provides a method and system for improving text content recognition and classification and computer equipment. The method comprises the steps of receiving an input text of a user; segmenting the input text through a word segmentation device to obtain a text word segmentation set; based on the data augmentation strategy and the text word segmentation set, expanding the keyword list and the semantic detection model to obtain training data of an expanded-writing keyword list and an expanded semantic model; judging whether the text word segmentation set is matched with the word list of the expanding-writing keywords or not; if not, processing the text word segmentation set into a model format to obtain a processed text; inputting the processed text into a pre-training semantic detection model, and judging whether the processed text is a non-compliant text or a sensitive text or not through a confidence coefficient threshold value; and if not, outputting the processed text. According to the method, through a hierarchical processing architecture of keyword matching and a semantic detection model, in combination with data augmentation and edge end lightweight deployment, the requirements of real-time processing and high accuracy are balanced, the real-time performance of text recognition and processing is improved, and the method is suitable for various application scenes.
Owner:SHENZHEN ZHUOYUE ZHIYUN TECHNOLOGY CO LTD

Digital human interaction method and system, storage medium and program product

The invention relates to the technical field of artificial intelligence, and discloses a digital human interaction method and system, a storage medium and a program product, which can associate a user with an interaction process by adding a target session identifier to a first voice request, ensure context coherence of multiple rounds of conversations, and realize full duplex interaction. Furthermore, the voice signal is converted into the processable target text, the command task in the target text is identified, and the first command task text is generated, so that the intention of the user can be accurately identified, the task type can be distinguished, and the semantic understanding efficiency is improved. Furthermore, through parallel processing of text sentence segmentation and speech synthesis, the mechanical feeling of single word output is avoided, the waiting time of the user is shortened, and the interaction naturalness is improved. Furthermore, the target command task voice and the second command task text are sent to the client, so that bimodal synchronous output is realized, the interaction naturalness is improved, multi-scene adaptation is realized, and the user experience is improved.
Owner:SHANGHAI INVESTIGATION DESIGN & RES INST CO LTD

Knowledge base and knowledge graph dynamic access control method and system based on AI classification and grading

The invention discloses a knowledge base and knowledge graph dynamic access control method and system based on AI classification and grading, and relates to the technical field of data security and access control. Automatically classifying and grading text segments or knowledge graph entities by using an AI model through a data preprocessing module, generating structured tags, and binding the structured tags with original data; after the data are associated and stored to the database through the structured storage module, the authority management module maintains the accessible classification and grade range of the user; and the query filtering module calls the user permission and dynamically filters the retrieval result when the user retrieves, and only returns the authorized content. According to the method, automatic and fine-grained dynamic access control based on data content semantics is realized, the problems of dependence on manual labeling, coarse access control granularity and insufficient flexibility in the prior art are solved, data security and retrieval efficiency are considered, and the method is adaptive to multiple coincidence rules and standards.
Owner:DANGKANG DATA INTELLIGENCE TECHNOLOGY (GUANGDONG) CO LTD

Document retrieval and conflict detection system based on retrieval enhancement generation

The invention discloses a document retrieval and conflict detection system based on retrieval enhancement generation, and the system comprises a document loading module which is used for analyzing a multi-format document and extracting text content; the text segmentation module is used for dividing the text content into semantic segments; the vector index module is used for converting the semantic fragments into vectors and constructing indexes; the retrieval module is used for receiving user query and retrieving related semantic fragments; the conflict detection module is used for inputting the semantic fragments queried and retrieved by the user into a large language model for comparative analysis so as to identify logic or fact conflicts; the streaming output module is used for returning an intermediate reasoning process and a final conclusion of the large language model in real time; and the batch processing module is used for automatically processing the documents of the specified directory and batch query and generating a structured result file. According to the method, the system transparency and the user credibility are improved, and the robustness and efficiency in practical application are improved.
Owner:BEIJING TECH & BUSINESS UNIV

Streaming voice service concurrent processing method based on soft switch

The invention relates to a streaming voice service concurrent processing method based on soft switching, and the method comprises the steps: building and maintaining a global voice recognition channel with an ASR server, continuously forwarding a user voice stream to the ASR server during a whole session period, and receiving a recognition result; processing the recognition result by adopting a finite-state machine and a large language model to generate a response text for multiple rounds of dialogues; and concurrently executing TTS processing while the global speech recognition channel is continuously operated: segmenting a response text into a plurality of segments according to semantics to perform streaming speech synthesis, and immediately sending a pseudo response after receiving first speech data so as to trigger a soft switch system to pre-load a playing buffer area in advance. Through global reception and concurrent processing, smooth multi-round dialogues which can be interrupted at any time are realized, and the interaction delay is reduced from a second level to a millisecond level by using a text fragmentation and early media triggering technology, so that the user experience and the interaction efficiency of an intelligent call center are remarkably improved.
Owner:上海井星信息科技有限公司

Method and system for generating community policing risk compensation measures based on graph search enhancement generation

The application relates to the technical field of public security risks, and provides a community public security risk compensation measure generation method and system based on atlas retrieval enhancement generation. The method performs text segmentation and vectorization on obtained domain knowledge data in the field of public security risks, generates knowledge slices in the field of public security risks and corresponding knowledge vectors, and stores the knowledge slices and the knowledge vectors into a vector database; meanwhile, a domain knowledge atlas in the field of community public security risks is constructed, and an index relationship between the domain knowledge atlas and the knowledge vectors is stored in the vector database, so as to construct a domain knowledge base in the field of public security risks; then, a user risk query is analyzed through a large language model, risk elements related to the user risk query are searched in the domain knowledge atlas, and knowledge slices corresponding to the risk elements are searched in the domain knowledge base according to the index relationship, so as to generate public security risk compensation measures corresponding to the user risk query.
Owner:CITIC GUOAN INFORMATION TECH +1

Long text generation method based on large language model

The invention provides a long text generation method based on a large language model, and relates to the technical field of text generation processing. S2, preprocessing a long text, and cooperatively generating the long text based on a large language model, including S2.1, text segmentation, S2.2, keyword extraction and S2.3, meta-information labeling; step S3, core generation: processing the preprocessed text information based on a large language model, including S3.1, hierarchical generation, S3.2, memory mechanism and S3.3, condition control; step S4, carrying out back-end processing; according to the method, the long text is segmented into semantically coherent paragraphs or sentences through the text segmentation algorithm, so that the LLM can more effectively process each segmented text unit, the problems of logic inconsistency and chaos caused by limited context windows are avoided, and the segmentation strategy not only improves the long text processing capability of the LLM, but also improves the segmentation efficiency of the LLM. And the overall quality and continuity of the generated text are ensured.
Owner:INFORMATION & COMMUNICATION BRANCH STATE GRID JIBEI ELECTRIC POWER CO LTD +3

Text style conversion control method and device, equipment and medium

The invention relates to the technical field of natural language processing, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a text style conversion control method, device and equipment and a medium. Performing text word segmentation on the source text and the target style description text to obtain a corresponding source text sequence and a corresponding target style sequence; performing style feature extraction on the source text sequence and the target style sequence to obtain corresponding source text style features and target style features; determining a target style feature conversion parameter according to the source text style feature and the target style feature; performing text coding on the source text sequence to obtain a source coded text sequence; and performing style conversion on the source coding text sequence according to the target style feature conversion parameter to obtain a target style text. The text style conversion efficiency and style control granularity can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-agent cooperation-based long text news event element extraction method and system

This invention relates to the field of information extraction technology, and discloses a method and system for extracting elements from long-text news events based on multi-agent collaboration. The method includes the following steps: constructing a large language model multi-agent system architecture; a central coordinating agent receiving unstructured news text; a text segmentation agent outputting logical units and their time descriptions; an event extraction agent extracting the event subject, action, and object, and completing time and entity role information; a scene matching agent determining the fine-grained news scene category to which it belongs; an information extraction agent generating structured event records; and a central coordinating agent performing quality review and anomaly handling on each step of the output to ensure the integrity and consistency of the output data. This invention, through relationship building and contextual analysis agents, can construct static relationships between event subjects based on structured event information, and summarize event situation changes based on a timeline, improving the feasibility and visualization of the system's news event analysis.
Owner:CENT SOUTH UNIV

Commodity label generation method and device for e-commerce scene, and storage medium

The invention relates to an e-commerce scene-oriented commodity label generation method and device, and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the steps of obtaining native text information and an image set of a target commodity; performing information identification and extraction on the image set, converting related visual data into structured text description data, and fusing the structured text description data into a to-be-processed text corpus with multi-dimensional information; dynamically generating a theme label structure for the target commodity; performing text segmentation on the to-be-processed text corpus according to the topic tag structure to obtain grouped texts corresponding to different topic tags, and feeding back each grouped text to a corresponding topic expert sub-model to perform candidate tag extraction; and semantic fusion is carried out, and a fused label set is used as a final label of the target commodity. According to the method, the integrity, accuracy and robustness of commodity label generation in an e-commerce scene are remarkably improved, and the limitation of dependence on manpower or a fixed label library is reduced.
Owner:HANGZHOU NO TABLE ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD