Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

150 results about "Semantic integrity" patented technology

Semantic integrity. Semantic integrity ensures that data entered into a row reflects an allowable value for that row. The value must be within the domain, or allowable set of values, for that column. For example, the quantity column of the items table permits only numbers.

Intelligent agent visual language navigation method and system based on task completion prediction

The invention provides an agent visual language navigation method and system based on task completion prediction. The method comprises the step of constructing a dual-drive structure composed of a self-adaptive mixed pooling mechanism and a task completion analysis module. Firstly, in the visual information processing process, a dynamic weight distribution strategy is adopted to carry out multi-scale adaptive mixed pooling on panoramic features, so that the fusion effect of local and global semantic information is optimized, and the retention capability and semantic integrity of navigation historical information in a dynamic topological map are improved. And then, inspired by a human navigation cognitive behavior mechanism, a task completion analysis module is designed and introduced, and the task execution progress is dynamically estimated based on the recognition condition of a key landmark in a navigation path, so that an intelligent agent is driven to preferentially select a key path node and invalid exploration is reduced. And finally, realizing efficient understanding and execution of the natural language instruction by the intelligent agent through a multi-round cyclic cross-modal reasoning and action prediction mechanism.
Owner:FUZHOU UNIV

Knowledge question-answering method and system based on topic knowledge graph retrieval enhancement

The invention discloses a knowledge question-answering method and system based on topic knowledge graph retrieval enhancement, and the method comprises the steps: firstly extracting a local topic represented in a triple form based on an original document through employing a large language model, carrying out the clustering, and generating a global topic triple set representing the global perspective of the whole document; secondly, on the basis of the global topic triple set, topic-guided entity and relation extraction is adopted, and a mixed knowledge graph is constructed; secondly, providing a semantic perception personalized PageRank algorithm, matching query semantics with semantics of edges in the mixed knowledge graph, and dynamically adjusting the weight of score propagation between nodes; and finally, designing a three-level progressive retrieval mechanism, retrieving multi-level information related to user query from the mixed knowledge graph, and inputting the multi-level information into the large language model to generate a final answer. According to the method, the semantic integrity and retrieval precision of the knowledge graph are remarkably improved, and the accuracy, comprehensiveness and enabling performance of generated answers are ensured.
Owner:HANGZHOU DIANZI UNIV

Text abstract generation method and system based on sparse attention acceleration

The invention discloses a text abstract generation method and system based on sparse attention acceleration, and the method comprises the steps: reading long text data, carrying out word segmentation and embedded coding processing, extracting a sequence feature vector, and mapping the sequence feature vector into a query matrix Q, a key matrix K and a value matrix V; constructing an abstract generation network, wherein the abstract generation network comprises a sparse attention calculation module, a feedforward calculation module, a prediction head module and a key value cache module; inputting the query matrix Q, the key matrix K and the value matrix V into an abstract generation network, passing through a backbone network formed by stacking a sparse attention calculation module and a feedforward calculation module for multiple times, processing by a prediction header module, and caching a historical decoding state in real time through a key value caching module to obtain an initial text feature vector; and based on the initial text feature vector, executing an autoregressive decoding process through an abstract generation network, and outputting a final abstract result. According to the method, the decoding process can be accelerated, and the semantic integrity and coherence of the generated abstract are ensured.
Owner:ZHEJIANG UNIV

Directory perception-based long document knowledge base construction method and program product

The invention discloses a long document knowledge base construction method based on directory perception and a program product, and belongs to the field of artificial intelligence and natural language processing. According to the scheme, an original long document is sequentially subjected to preprocessing, directory structure analysis, mixed blocking, double-tag generation, tag intelligent optimization, vectorization and meta-information mounting, and finally automatic construction and high-quality retrieval enhancement generation of a knowledge base are achieved. According to the method, the semantic integrity is guaranteed by fully utilizing the perceptual ability of the directory structure, the label quality and the retrieval efficiency are improved by combining a double-label system and intelligent optimization circulation, and the construction efficiency, the retrieval accuracy and the result traceability of the knowledge base are remarkably improved; the method is suitable for intelligent processing and application of long complex structure documents such as academic specialities, technical documents, policies and regulations and the like.
Owner:SOUTHEAST UNIV

Intelligent document information real-time retrieval method and system based on RAG technology

The invention relates to an intelligent document information real-time retrieval method and system based on the RAG technology, and the method comprises the steps: analyzing documents of various formats, extracting a text, maintaining the content continuity through adaptive semantic partitioning processing, and building a character-level position index at the same time; after vectorizing the text blocks, generating a plurality of rewriting queries for the original query; respectively carrying out mixed retrieval (combining keywords and semantic retrieval) for each rewriting query, and fusing and reordering results to obtain candidate text blocks; after correlation filtering, inputting a large language model according to correlation to generate an answer, and if the result is negative, triggering secondary retrieval and reordering; and finally, outputting a structured answer containing position information and supporting front-end visualization. According to the method, the semantic integrity maintenance, the multi-format document processing efficiency and the key information retrieval accuracy are effectively improved.
Owner:ECCOM NETWORK SYST CO LTD

Large language model multi-source inference mapping knowledge domain question and answer method and system

The invention discloses a large language model multi-source inference knowledge graph question answering method and system, and the method comprises the steps: extracting entities in a question, linking the entities with knowledge graph entities, and determining a subject entity set; searching paths in the knowledge graph by adopting a graph constraint reasoning method, and obtaining candidate answers and reasoning paths; detecting incomplete path coverage or answer dimension deficiency by using a large language model, triggering question decomposition to generate logically complementary sub-questions, and obtaining answers; establishing an inference source coordination layer, a coordination graph constraint inference source, a planning-retrieval inference source and a sub-problem inference source for multi-source evidence; and standardizing the output of each inference source into a structured evidence triple, sorting and selecting the first K evidences according to relevance, and generating a final answer by utilizing a large language model to carry out inductive inference. According to the method, reasoning can be dynamically adjusted, multi-source evidences can be effectively aggregated, logic consistency and semantic integrity are guaranteed, and the integrity and accuracy of knowledge graph multi-hop questions and answers are improved.
Owner:BEIFANG UNIV OF NATITIES

Medical document slicing and retrieval method and device, equipment and storage medium

The invention discloses a medical document slicing and retrieval method and device, equipment and a storage medium, and the method comprises the steps: generating a high-fidelity semantic summary text and a child node of a structured keyword set through a father node containing an original complete paragraph text, and storing the child node to a vector database; storing the semantic summary text and the structured keyword set into a keyword database; and when a user query instruction is received, performing dual-channel retrieval through the vector database and the keyword database, obtaining related target child nodes, extracting target original contents of a target father node from the target child nodes, transmitting the target original contents to the large language model, generating a final answer, and sending the final answer to the user. According to the method, the relevance and accuracy of retrieval results can be remarkably improved, instant and accurate decision support basis is provided for doctors, semantic integrity can be reserved, efficient multi-modal retrieval can be achieved, and the speed and efficiency of medical document slicing and retrieval are improved.
Owner:XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV

Robust noise reduction processing method and system for sound wave signal self-supervised learning enhancement

ActiveCN121415799ASpeech analysisPhysical realisationTime domainProbability propagation
The invention provides a sound wave signal self-supervised learning enhanced robust noise reduction processing method and system, and relates to the technical field of signal processing, and the method comprises the steps: carrying out the feature enhancement of an initial time-frequency representation through a dynamic adaptive mask strategy, constructing a self-supervised reconstruction task based on the mask time-frequency representation, separating noise and signal subspaces in a semantic manifold space, and carrying out the self-supervised learning enhanced robust noise reduction. And establishing a probability propagation network in combination with time domain continuity characteristics to model a local dependency relationship, and finally generating a noise reduction weight and realizing semantic fidelity optimization. The method can effectively improve the noise reduction effect and semantic integrity of sound wave signals in a noise complex environment.
Owner:BEIJING GUANYU INFORMATION TECHNOLOGY CO LTD

Multi-language text adaptive configuration method and electronic equipment

The invention discloses a multi-language text self-adaptive configuration method and electronic equipment, and relates to the technical field of computers, and the method comprises the following steps: obtaining a source text set in response to a translation instruction received when a page runs, and translating the source text set according to a target language identifier and a constraint strategy to obtain a translated text set; according to a constraint strategy, carrying out adaptability detection on each translation in the translation set to obtain an adaptive set and a non-adaptive set; performing semantic rewriting on each non-adaptive translation in the non-adaptive set to obtain a target candidate set; and aggregating the adaptation set and the target candidate set in the same rendering frame, and rendering and displaying the adaptation set and the target candidate set in batches, so that the problems of lack of multi-language dynamic adaptation capability and insufficient semantic equivalent compression are solved, the accumulated layout offset is remarkably reduced on the premise of ensuring semantic integrity and readability, and the layout efficiency is improved. And the page stability and the user experience are improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Image fusion method and system based on multi-semantic guidance and mixed experts

The invention discloses an image fusion method and system based on multi-semantic guidance and mixed experts. The problems that in the prior art, robustness is insufficient, and visual fidelity and semantic integrity are difficult to balance are effectively solved. According to the method, infrared and visible light images to be fused and task identifiers are obtained, and firstly, a CLIP network is utilized to extract high-level semantic features as guide vectors; and then inputting the image and the guide vector into a pre-trained hybrid expert (MoE) fusion network. According to the network, intermediate features are extracted through an encoder, a gating network dynamically activates part of expert subnets according to task identifiers and semantic vectors and calculates routing weights, adaptive nonlinear transformation and weighted fusion are carried out on the features, and finally a high-quality fusion image is reconstructed through a decoder. The method can adapt to different task requirements, and the calculation efficiency is remarkably improved while the image fusion quality and the semantic consistency are improved.
Owner:XIDIAN UNIV

Multi-industry information system integrated heterogeneous data fusion processing platform

The invention relates to the technical field of heterogeneous data fusion processing, in particular to a heterogeneous data fusion processing platform for multi-industry information system integration, which comprises an industry knowledge graph module for acquiring business data of each industry, extracting a core entity relationship of each industry and constructing a field knowledge graph of at least two industries; when the system is used, the limitation of fixed weighted average is broken through by constructing a lightweight domain knowledge graph for each industry, extracting a core entity relationship and dynamically adjusting the fusion weight of each data feature, so that the purpose of dynamically adapting to the semantic feature difference of different industries is achieved; the method can be applied to cross-department data fusion of smart cities conveniently, semantic conflict intensity is accurately quantified by using a dual-channel neural network, adaptive strategy grading processing is combined, efficient resolution is performed through unit conversion during low conflicts, semantic integrity is guaranteed by using field isolation fusion during high conflicts, and conflict resolution efficiency and accuracy can be improved.
Owner:HUAIAN XINGMINGCHUANG INFORMATION TECHNOLOGY CO LTD

Real-time audio-video synchronous generation method and system based on semantic analysis

The invention relates to the technical field of audio and video generation, in particular to a real-time audio and video synchronous generation method and system based on semantic parse, and the system and method respectively decompose texts, audios and videos into minimum semantic units, and combine with a pre-training model to ensure that each modal unit is complete in semantics and accurate in granularity. And meanwhile, a cross-modal context consistency coefficient is introduced to filter pseudo associations with similar semantics but irrelevant scenes, and a matching uniqueness punishment mechanism is matched to avoid generation of conflicts, so that the problems that semantic associations are fuzzy and contents deviate from requirements in traditional video and audio generation are effectively solved. Video and audio synthesis based on an optimal text-audio and video semantic unit matching combination output by game optimization and core strategy parameters is a key advantage of guaranteeing generation quality. The accurate corresponding relation between the text and the audio and video unit is defined through the optimal matching combination; the core strategy parameters provide a dynamic adaptation basis for the synthesis process, and audio and video generation modules can be guided to adjust quality parameters and resource allocation according to scene requirements.
Owner:YONGBAO JIAFU (SHANGHAI) IND CO LTD

Long text processing method and device based on large model, electronic equipment and medium

The invention discloses a long text processing method and device based on a large model, electronic equipment and a medium, relates to the technical field of text processing, and solves the technical problems of lack of targeted processing rules, lack of systematic methods for correction and global integration of blocked overlapped contents, and easy occurrence of information inconsistency or omission. According to the method, through dual classification of text lengths and types and matching of differentiated partitioning strategies, semantic integrity is guaranteed, processing efficiency is improved, five conflict types including facts, logic, semantics, anaphora and ranges are covered, accurate positioning of multi-dimensional conflicts is achieved in combination with professional tools, limitation of single-dimensional analysis is avoided, and the method is suitable for large-scale popularization and application. Differentiated processing rules are designed for different conflict types, and the accuracy of conflict correction is ensured; and meanwhile, through overlapped content correction and progressive integration, global information consistency and completeness are achieved, an integration result is verified through a secondary processing mechanism, and the consistency and completeness of long text processing are further guaranteed.
Owner:BEIJING SHIJI INTELLIGENT TECHNOLOGY CO LTD

Large model information extraction and structure restoration system for long text document

The invention belongs to the technical field of intelligent document processing, and particularly relates to a large model information extraction and structure restoration system for a long text document. The system comprises an extraction point defining and text preprocessing module which is used for carrying out deep preprocessing and analysis on an input unstructured document, extracting text and coordinate information and identifying and protecting a special structure; meanwhile, an intelligent dynamic blocking strategy based on semantic boundaries is adopted, a dynamic overlapping mechanism is combined, and an input document is divided into text blocks keeping semantic integrity; the high-concurrency processing and scheduling control module is used for processing the text blocks in batches, calling a large language model to carry out distributed reasoning and outputting a dispersion result; and the extraction reasoning and result fusion module is used for performing semantic deduplication, entity alignment and confidence fusion on the dispersion result returned by the large language model to generate globally consistent structured output. The method has the advantages of being high in precision, high in speed, controllable in cost and high in robustness.
Owner:浙江实在智能科技有限公司

Method for compressing thinking chain of reasoning large model

The invention discloses an inference large model thinking chain compression method, which comprises the following steps of: firstly, generating answer sets with different detailed degrees by utilizing multiple rounds of sampling of a basic large model, and adaptively selecting an inference chain length by adopting a dynamic quantile algorithm based on task difficulty; secondly, performing diversity rewriting and compression on the reasoning step through KL divergence constraint, and generating the shortest expression on the premise of ensuring semantic consistency; constructing positive and negative samples to guide the model to learn simple expression, and training by adopting a composite loss function including supervised learning and length perception preference optimization; according to the method, external annotation data or a teacher model is not needed, adaptive matching of the reasoning depth and the problem difficulty can be achieved, the semantic integrity is guaranteed, meanwhile, the reasoning efficiency is remarkably improved, high expandability and good cross-task migration ability are achieved, and the method is particularly suitable for large-model lightweight deployment in a low-computing-power environment.
Owner:ZHEJIANG UNIV

Internet webpage content feature extraction method based on artificial intelligence

The invention discloses an Internet webpage content feature extraction method based on artificial intelligence, which comprises the following steps: S1, acquiring a webpage HTML source file and a rendering image, and preprocessing the webpage HTML source file and the rendering image; s2, performing semantic classification on DOM nodes, encoding the DOM nodes into three types of identifiers, and constructing a node label sequence; s3, performing time sequence synchronization on the node attribute vector, the node tag sequence and the visual area set, and performing block-level slicing; s4, inputting the block-level slices into a gated recursive attention network, extracting multi-modal joint representation, and executing attention aggregation; s5, performing structure alignment on the feature fusion sequence, calculating a cross-node consistency distance, and screening a target slice set with high structure cohesion; and S6, mapping the target slice set to a content feature space, and generating a content feature tag set. According to the method, the structural accuracy, the semantic integrity and the multi-modal fusion precision of webpage content extraction are improved.
Owner:NANJING YUANPENG SOFTWARE TECHNOLOGY CO LTD

Multi-mode voice interaction method and device, intelligent equipment and readable storage medium

The invention provides a multi-mode voice interaction method and device and intelligent equipment, is suitable for the technical field of intelligent voice interaction, is applied to the intelligent equipment, and comprises the following steps: in response to a detected voice activity, extracting first voice data in the voice activity, and obtaining video data shot synchronously with the first voice data, a plurality of different users are shot in the video data. And screening out a target user sending the first voice data from the video data. And performing semantic integrity analysis on the text content corresponding to the first voice data. And when the semantic integrity analysis result is that the semantics is incomplete, continuously acquiring second voice data of the target user for a plurality of times, and generating a corresponding target statement with complete semantics after the first voice data and the second voice data are combined. Generating reply data according to the target statement, and outputting the reply data through voice. According to the embodiment of the invention, accurate, coherent and real-time voice interaction with the user needing interaction can be realized.
Owner:浙江人形机器人创新中心有限公司

Document processing method and related equipment

The invention discloses a document processing method and related equipment, and belongs to the technical field of data processing.The method comprises the steps that in response to a document input instruction, a to-be-processed target document is obtained; performing structure analysis on a to-be-processed target document to generate a directory tree of the target document; calling a dynamic recursive slicing algorithm, traversing and analyzing the directory tree of the target document, and generating a plurality of text slices with dynamic lengths; and vectorizing the plurality of text slices with the dynamic lengths, and uploading and storing the vectorized text slices to a database of a retrieval enhancement generation system. By constructing the directory tree and traversing the directory tree by adopting the dynamic recursive slicing algorithm, the document structure can be more accurately understood, the semantic integrity of the document slices is improved, and the retrieval accuracy and the answer generation quality of the retrieval enhancement generation system are further effectively improved.
Owner:E-SURFING DIGITAL LIFE TECH CO LTD

Code summary generation method, model, device, medium and apparatus

This application provides a code digest generation method, model, apparatus, medium, and device, relating to the field of artificial intelligence technology. The method includes: determining an application programming interface (API) information graph corresponding to the code to be processed, wherein the node types in the API information graph include: a first type, and at least one of a second type and a third type, wherein the associated information of the first type of nodes is the application programming interface, the associated information of the second type of nodes is variables, and the associated information of the third type of nodes is API description information; encoding the API information graph to obtain API encoding features of the code to be processed; and determining a code digest of the code to be processed based on the API encoding features and target features of the code to be processed, wherein the target features include: source code encoding features corresponding to the code to be processed, or source code encoding features and AST encoding features corresponding to the code to be processed. This application embodiment can improve the semantic integrity of the code digest.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1

Document processing method, server, storage medium and program product

The invention provides a document processing method, a server, a storage medium and a program product. According to the method, a document hierarchy tree of a document is obtained, and the document hierarchy tree comprises title nodes corresponding to hierarchy titles in the document and text nodes corresponding to text segments in the document and comprises hierarchy structure information of the document; according to a bottom-up merging strategy that brother nodes are preferentially merged and then father nodes are merged, text nodes and title nodes in the document hierarchical tree are grouped and merged to obtain text block nodes, and texts corresponding to the text block nodes serve as text blocks of the document; according to the method, texts corresponding to brother nodes with closer relationships on a hierarchical structure can be preferentially merged into the same text block, and then texts corresponding to father nodes with relatively close relationships are sequentially merged from bottom to top, so that the semantic integrity, relevance and cohesion of the text block are improved; therefore, the accuracy of retrieving and enhancing the recalled text blocks is improved, and the response quality of human-computer interaction is effectively improved.
Owner:ALIBABA CLOUD COMPUTING CO LTD

Speech translation method, system, and storage medium

The application discloses a speech translation method and system and a storage medium, relates to the technical field of speech processing, and comprises the following steps: acquiring a real-time audio stream of source audio, and incrementally generating source language text corresponding to the real-time audio stream; detecting the semantic integrity of the generated source language text in real time; in the case where the semantic integrity is greater than or equal to the integrity threshold value corresponding to a minimum semantic unit, taking the generated source language text as a source language text block; determining target translation corresponding to the source language text block, and outputting the target translation. After the source language text is incrementally generated, the translation of the source language text block and the output are driven based on the semantic integrity, online segmentation translation of long sentences is realized, the translation waiting time is shortened, and the target translation and the original sound are synchronized.
Owner:ZHUHAI MOJIE TECH CO LTD

Business data automatic vectorization and semantic retrieval method and system

The invention relates to the technical field of semantic retrieval, and discloses an automatic vectorization and semantic retrieval method for business data, which comprises the following steps of: data change capture: sensing newly adding, modifying or deleting operation of data in a business system, extracting business records and binding record identifiers; multi-dimensional quality evaluation: carrying out five-dimensional quality evaluation of integrity, consistency, timeliness, information entropy and security on the extracted service data through an MDQA algorithm, and only allowing the quality standard data to enter a subsequent process; according to the method, through set multi-dimensional quality evaluation intelligent semantic segmentation and incremental vector updating, the retrieval quality is remarkably improved, 22% of low-quality data is filtered through the MDQA algorithm, the retrieval precision is improved by 19.9%, and the user satisfaction degree is improved by 28.1%; the SBAS algorithm ensures segmentation semantic integrity, the recall rate is improved by 25.8%, and retrieval deviation caused by context loss is avoided.
Owner:SUZHOU RUIYING INTELLIGENT TECH CO LTD

PDF (Portable Document Format) self-adaptive partitioning method, device, equipment, medium and product

The invention discloses a PDF self-adaptive partitioning method and device, equipment, a medium and a product, and relates to the technical field of data processing, and the method comprises the steps: carrying out the structure reconstruction of an analyzed original PDF document flow, and obtaining a reconstructed PDF document flow; identifying special semantic units, titles and guiding keywords of the reconstructed PDF document flow, integrally replacing the identified special semantic units, judging the continuity of the identified titles, determining segmentation priorities of the titles according to judgment results, and performing segmentation on the titles according to the segmentation priorities. Carrying out semantic binding on the guiding keyword meeting the requirement and the subsequent text to obtain a to-be-segmented PDF document stream; according to the method, the to-be-segmented PDF document stream is subjected to self-adaptive segmentation according to the segmentation priority by adopting the preset multi-level separators, the preliminary segmentation result is obtained, and compared with fixed-length segmentation, the coherence and semantic integrity of the document after PDF segmentation can be guaranteed.
Owner:BEIJING WISDOM TOOTH TECH CONSULTING CO LTD +1

Table metadata processing method and system, electronic equipment and storage medium

The invention relates to the technical field of computers, and discloses a table metadata processing method and system, electronic equipment and a storage medium, and the method comprises the following steps: carrying out regional semantic segmentation through weighted fusion based on frame geometric features, character coordinates and language features of a table; constructing a structural feature vector based on the row number, the column number and the row list header cell style of the table, performing geometric semantic integrity reconstruction, executing element standardization and semantic extraction, and performing geometric semantic integrity reconstruction on the basis of a region coding library, a unit conversion chain and a knowledge element rule library by utilizing normalized index names, time and region information. And calculating a regional code, a standard measurement value, a knowledge element ID and a retrieval field, matching the normalized index with an existing index system, and outputting a structured record. According to the method, the requirements of high accuracy and traceable structuralization during data processing in a big data environment can be met.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

A data storage system and method

Embodiments of the present application provide a data storage system and method, the data storage system comprises a data input end, a master node, at least two slave nodes, a first database and a second database, the master node and the at least two slave nodes are communicatively connected, and each of the slave nodes is communicatively connected, the master node is used for reading initial data from the data input end and sending the initial data to the at least two slave nodes, the at least two slave nodes are used for calculating the initial data according to a preset data processing flow to obtain device data and metadata, the master node is further used for receiving the metadata and storing the metadata in a mode into a HIVE table of the first database, and the at least two slave nodes are further used for storing the device data into the second database, the metadata calculated is stored in a mode, the problems of low data storage efficiency and high cost caused by using a JSON format for storage are avoided, and the semantic integrity of the initial data is ensured, and the readability of the data is not affected.
Owner:ROOTCLOUD TECH CO LTD

Speech interruption decision method and system based on multi-granularity semantic completeness prediction

PendingCN122313981ASemantic featureData mining
This invention discloses a speech interruption decision-making method and system based on multi-granularity semantic integrity prediction, belonging to the field of speech interaction technology. The method includes: extracting multi-granularity semantic features and combining them into a current multi-granularity semantic feature vector; calculating the offset of the current multi-granularity semantic feature vector with a standard semantic pause model to obtain a semantic offset feature vector; performing similarity matching and weighted voting between the semantic offset feature vector and a historical interruption decision case library to obtain the current predicted interruption confidence; and generating an interruption command when the current predicted interruption confidence exceeds a dynamic threshold. This invention solves the problems of existing technologies where speech interruption decisions rely on simple energy thresholds or semantic integrity probabilities, lack utilization of user expression habits and contextual experience, and have low decision accuracy by introducing a standard semantic pause model and a historical case analogy reasoning mechanism, thus achieving more intelligent and accurate speech interruption judgment.
Owner:GUANGZHOU JIUSI INTELLIGENT TECH CO LTD

Subtitle intelligent batch translation method based on semantic perception

The invention provides a subtitle intelligent batch translation method based on semantic perception. The subtitle intelligent batch translation method comprises the steps that S1, a semantic boundary recognition algorithm with a multi-dimensional batch rule is constructed according to semantic perception logic; s2, identifying a segmentation boundary of the caption text by using a semantic boundary identification algorithm; s3, setting constraint conditions according to the content difficulty, the minimum threshold value, the maximum threshold value and the context semantic integrity of the segmentation boundary of the current batch; dynamically adjusting the segmentation boundary according to a constraint condition; s4, constructing a hierarchical Prompt which comprises a role constraint layer, a term constraint layer, a high-value context layer and a task layer; s5, guiding the LLM model to complete translation of the current batch according to the hierarchical Prompt, screening out high-value information in the current batch, and reserving the high-value information in the next batch; and S6, repeatedly executing the steps S2-S5 until translation of all batches is completed. According to the method, the effect of greatly improving the quality and efficiency of subtitle translation is achieved by accurately sensing the semantic boundary, dynamically and reasonably batching and optimizing translation output layer by layer.
Owner:ZHUHAI MAIYUE INFORMATION TECH

Conversation processing method and device and computer equipment

The invention provides a dialogue processing method and device and computer equipment, and the method comprises the steps: determining first information based on current input information; cutting the first information by using a cutting model to obtain second information, the second information representing key information of the first information, and the second information serving as input of a preset dialogue model to generate reply information corresponding to the current input information; wherein the cutting model is obtained by performing reinforcement learning training by using a preset multi-dimensional reward function, and the multi-dimensional reward function is used for updating parameters of the to-be-trained cutting model at least once. In this way, the first information is clipped through the clipping model trained by using a multi-dimensional reward function and a reinforcement learning feedback mechanism, so that the output second information keeps key information, and meanwhile, the semantic integrity and the length controllability are considered.
Owner:LENOVO (BEIJING) LTD

Retrieval enhanced question and answer method, device and equipment for semantic repair

The invention discloses a semantic repair retrieval enhanced question and answer method, device and equipment. And determining each candidate content combination based on each related content segment and a hierarchical logical relationship among the content segments in a pre-stored reference document, performing semantic repair on each related content segment based on the hierarchical logical relationship of the reference document, and ensuring that the content segments subjected to semantic repair have logical relevance. And semantic fracture or understanding deviation caused by fragmented information is avoided. The preset restoration condition comprises that the total length of the total content segments in the candidate content combination is smaller than a preset length threshold value, and the proportion of each related content segment in the candidate content combination in the total content segments reaches a preset proportion threshold value, and the content segments after semantic restoration can be screened through the preset restoration condition. The waste of computing resources caused by overlong full content fragments is avoided, and excessive interference of redundant information is avoided, so that the semantic integrity and the retrieval precision are balanced.
Owner:BEIJING UNISOUND INFORMATION TECH CO LTD

A method and system for intelligent structured processing of a regulation document, and a medium

The application discloses a kind of intelligent structured processing method and system of rules and regulations document, belong to artificial intelligence and natural language processing technical field.Along with the problem that the semantic of long regulation document in prior art is split due to the input length limit of large language model, output is unreliable, the application fuses the paragraph type by small model and semantic feature recognition, corrects label in combination with business rules, and aggregates semantic complete text fragment according to regulation level logic;The segment is constructed as guide prompt input large language model to generate structured result along with type label, and JSON format and field integrity check are set, and when check fails, it is returned to small model result.The method guarantees the integrity of clause-level semantics, improves the accuracy of key field extraction, and ensures system robustness.It is suitable for enterprise compliance management, regulation knowledge base construction and other scenarios, to realize the reliable conversion of unstructured regulation document to unified structured data.
Owner:YGSOFT INC