Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

171 results about "Semantic clustering" patented technology

Semantic clustering helps your company discover gaps in your content to enrich your customer’s experience. Inbenta’s Semantic Clustering groups semantically equivalent search queries — words, phrases and sentences — into clusters based on meaning. The higher the number of questions, words and phrases with a similar meaning, the greater the cluster.

SQL intelligent generation method and system for business query

The invention provides an intelligent SQL generation method and system for business query, and belongs to the technical field of artificial intelligence. Related data of query statements are acquired, and a corresponding query intention knowledge graph is constructed by semantic clustering; the method comprises the following steps: analyzing a historical SQL statement structure, extracting a natural language template and an SQL template, and expanding through a large language model to generate a feed-shot example set; and constructing a composite cue word template by combining task setting guidance, a feed-shot example and CoT chain thinking reasoning guidance. An intention completion module is arranged in a large language model, a natural language query statement of a user is combined with a composite cue word template context, a structured query statement is generated through entity recognition, semantic completion, parameter filling and fuzzy intention training, and the structured query statement is converted into a standard SQL statement through a knowledge graph and a template. According to the method, the use threshold of business personnel is remarkably reduced, and efficient conversion from natural language questions to SQL statements is realized.
Owner:国网福建省电力有限公司营销服务中心 +1

Data processing method and system for enterprise digital transformation platform

The embodiment of the invention provides a data processing method and system for an enterprise digital transformation platform, and belongs to the field of data processing. The method comprises the steps that structured field information from all heterogeneous data sources is acquired, and preprocessing operation is executed on the structured field information; constructing the processed structured field information into an embedded input sequence, splicing the embedded input sequence into a natural language fragment according to a preset template, and inputting the natural language fragment into a fine-tuned semantic coding model to obtain a corresponding semantic embedded vector; identifying similar field groups by adopting a clustering algorithm based on density or a hierarchical structure, and classifying each group of structured field information into a semantic cluster; and generating a corresponding standard field identifier for each semantic clustering cluster, and storing the generated standard field identifier in a standard field index database of the platform after digital transformation. According to the scheme, the field unified management and cross-system data alignment capability of the enterprise digital platform is remarkably enhanced.
Owner:YIBIN DIGITAL ECONOMY IND DEVELOPMENT CO LTD

Government affair work order intelligent processing method and system based on space-time semantic clustering and large language model

The invention relates to the field of government affair work order intelligent processing, in particular to a government affair work order intelligent processing method and system based on space-time semantic clustering and a large language model. According to the scheme, unified data feature modeling is conducted on a work order to be processed, an improved DBSCAN clustering algorithm is executed on the work order through a weighted space-time semantic three-dimensional distance measurement formula, and combined clustering of space, time and semantic features is achieved; calculating priority scores of the work orders, and dynamically allocating scheduling resources according to the clustering scale and the priority of the work orders; based on a retrieval enhancement generation technology of an RAG framework and an FAISS vector retrieval library, historical similar work orders are matched, a few-sample learning case is generated, and two sets of differential treatment schemes are generated by controlling temperature parameters of a large language model; visual display and interactive analysis of work order clustering are realized through an interactive GIS platform; and establishing a quality feedback closed loop of work order reconstruction. The method is suitable for intelligent government affair work order processing.
Owner:THE CHINESE UNIV OF HONG KONG (SHENZHEN) +2

Semantic analysis fused flow chart automatic layout method

The invention discloses an automatic flow chart layout method fusing semantic analysis, which relates to the technical field of automatic flow chart layout, and comprises the following steps of: inputting a node set with a business ring dependence condition into a time sequence conflict analysis engine, and combining a semantic vector and a time sequence vector to obtain a time sequence conflict analysis result; calculating a time sequence conflict index of the annular dependency set by adopting a weighted path consistency check algorithm so as to determine a time sequence conflict degree under the condition that the nodes have service annular dependency, and generating corresponding conflict description data; and inputting the conflict description data and the semantic vector into a conflict perception clustering optimizer, introducing a time sequence conflict penalty term into a clustering objective function, and performing cluster boundary adjustment on the annular dependency set through a spectral clustering algorithm so as to adjust a semantic clustering structure according to a determination result. According to the method, the problem that a semantic clustering structure cannot be optimized in combination with a time sequence conflict under business annular dependence is solved, and the effects of conflict accurate identification, clustering boundary dynamic adjustment and layout saliency enhancement are achieved.
Owner:XIAN XUNSHENG INFORMATION TECH CO LTD

Multi-agent traceable analysis method, device, equipment and medium

The invention relates to the technical field of data analysis, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a multi-agent traceable analysis method, device, equipment and medium, and the method comprises the steps: receiving a target theme and a data source list, collecting a multi-source document, and carrying out the preprocessing of the multi-source document to generate a preprocessing document set; configuring an analysis agent based on a semantic clustering result, and setting an analysis direction to form an analysis agent set; generating a structured note and index data table, and executing cross-document comparison to form an analysis output set; and receiving a feedback instruction to adjust the analysis agent set, triggering incremental processing to update the analysis output set, generating a theme research and judgment report, and keeping mapping consistency. According to the method, the multi-source document is fused through semantic clustering and a multi-agent cooperation mechanism, semantic association and traceable analysis are achieved, agent configuration is optimized in combination with interactive feedback, and the accuracy and the intelligent level of report generation are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent composition quality evaluation method and system based on large language model

The invention relates to the technical field of artificial intelligence in the education industry, in particular to an intelligent composition quality evaluation method and system based on a large language model, and the method comprises the steps: carrying out the text normalization and semantic unit segmentation of a composition, extracting a semantic vector through a first large language model in combination with a context enhancement strategy, and positioning a semantic fracture risk position; recognizing composition core elements through a second large language model, and mapping the composition core elements back to the semantic unit sequence; constructing a demonstration logic diagram, extracting a core demonstration path and abstracting the core demonstration path into a logic role topological graph; in combination with a pre-constructed writing specification knowledge graph, comparing structural compliance, connection strength and an expected support relationship, identifying and demonstrating logic defects, and generating a global deduction item list; semantic clustering is carried out on illegal items to form an error label set, comprehensive weight is calculated in combination with historical data of students, and core weak items are positioned; according to the application, the logic analysis depth of intelligent evaluation of the argument is remarkably improved, and the pertinence and practicability of teaching feedback are improved.
Owner:DALIAN HOUREN EDUCATION TECH CO LTD

Financial robot invoice element identification method based on semantic extraction

The invention discloses a financial robot invoice element identification method based on semantic extraction, and the method comprises the following steps: S1, obtaining an original invoice image, and carrying out the image preprocessing; s2, executing optical character recognition operation, and extracting invoice text information; s3, inputting a semantic potential model; s4, constructing a semantic kernel vector set according to preset invoice element categories; s5, generating a potential tensor field based on the semantic kernel vector set; s6, performing iterative semantic migration operation on the character units in the potential tensor field to form a semantic clustering region; s7, calculating a comprehensive confidence score, and outputting an invoice element recognition result; and S8, performing field legality verification on the invoice element identification result, and submitting the invoice element identification result to a financial robot system after verification is passed to drive related business processes. According to the method, semantic potential modeling and context coding technologies are fused, invoice elements are accurately extracted, and the method has the advantages of being clear in structure, high in robustness and high in adaptability.
Owner:LIANYUNGANG GUOTU INFORMATION TECHNOLOGY CO LTD

Financial user portrait analysis method and system based on knowledge graph

The invention discloses a financial user portrait analysis method and system based on a knowledge graph, and the method comprises the steps: recognizing a financial entity through an FNER algorithm, carrying out the relation embedding through employing a TransFinE algorithm, integrating the time weight attenuation and risk propagation constraints, constructing a financial knowledge graph, calculating the user influence distribution through employing a FinRank algorithm, and carrying out the analysis of the financial user portrait. And finally carrying out semantic clustering and grouping to obtain a user portrait classification result. The technical problems that an existing financial user portrait analysis method cannot effectively process a multi-dimensional semantic association relationship, lacks a user network influence quantification mechanism and neglects time sequence evolution characteristics are solved.
Owner:IND & COMMERCIAL BANK OF CHINA LTD

Multi-dimensional confidence fusion large language model uncertainty evaluation method and system

The invention belongs to the technical field of natural language processing, provides a multi-dimensional confidence fusion large language model uncertainty evaluation method and system, designs a multi-dimensional confidence modeling mechanism, and can perform multi-dimensional confidence modeling on the basis of internal information in a large language model generation process without depending on an external knowledge base. And the reliability of the output content is effectively judged. According to the system, on the basis of a semantic clustering mechanism, a Token-level multi-dimensional confidence modeling method is innovatively introduced, a scoring system fusing factors such as probability centrality, context disturbance sensitivity, generation consistency and language rationality is constructed, the confidence structure of each part of content in model output can be evaluated from the Token level, and the evaluation efficiency is improved. And the discrimination capability of the system on the uncertainty difference in the generation process is obviously improved.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +3

Index optimization and compression storage system and method for large-scale literature set

The invention discloses an index optimization and compression storage system and method for a large-scale literature set, and the method comprises the following steps: S1, collecting and preprocessing literature data, and generating a standardized text data set; s2, carrying out keyword semantic vector coding, and constructing a keyword semantic vector matrix; s3, constructing an initial Gaussian mixture model to obtain a clustering center, a covariance matrix and a weight; s4, introducing a sea elephant optimization algorithm to optimize clustering parameters, and outputting an optimal clustering result; s5, constructing a semantic clustering structure, and generating an index tree structure; s6, performing bitmap compression and inverted coding, and constructing an index table supporting Boolean logic; and S7, dynamically accessing the newly added literature, and completing incremental updating of the index structure. The method is used for improving the index construction efficiency and the storage compression rate of a large-scale literature set, and efficient and semantic literature retrieval service capable of being incrementally updated is achieved.
Owner:CENTRAL COMPILATION & TRANSLATION PRESS CO LTD

News viewpoint evolution trend tracking system based on artificial intelligence

The invention relates to the technical field of artificial intelligence, in particular to a news viewpoint evolution trend tracking system based on artificial intelligence, which comprises a collector, an extractor, an aggregator, a modeler, a detector and a presenter. A disturbance sample generation and consistency detection mechanism is combined, so that the clustering process is more robust, and the recognition precision of the viewpoint cluster semantic center is remarkably improved; according to the method, a black hole detection mechanism based on density estimation and multi-point similarity joint modeling, an anti-fact intervention mechanism and a dynamic graph density-structure joint estimation model are adopted, a node disturbance situation can be constructed, the influence of the node disturbance situation on a semantic structure can be inferred, and through fusion of semantic consistency, density offset and energy anti-fact scoring, the dynamic graph density-structure joint estimation model is obtained. High-confidence screening of abnormal nodes is realized, and the purity of a viewpoint atlas structure and the reliability of semantic evolution modeling are effectively improved.
Owner:HEBEI UNIVERSITY

Electronic component online sales data management and maintenance system

The invention relates to the technical field of sales data management, and discloses an electronic component online sales data management and maintenance system, which comprises a data acquisition and processing module, a user demand matching module, a data verification module, an inventory optimization module and the like. The data acquisition and processing module acquires multi-dimensional sales data, and performs natural language processing and dynamic knowledge graph construction, processing and data association; the user demand matching module generates an initial matching result based on a collaborative filtering algorithm and semantic clustering; the data verification module verifies the data by using a Hash algorithm and a block chain network; and the inventory optimization module optimizes the inventory according to the timeliness weight factor and a reinforcement learning algorithm. In addition, the system is further provided with an abnormal transaction detection module, a multi-source data integration module and an interactive query optimization module which are respectively used for detecting abnormal transactions, integrating multi-source data and optimizing query. The system effectively solves the problem of electronic component online sales data management, and improves sales efficiency and user experience.
Owner:SHENZHEN KELLYXUN TECHNOLOGY CO LTD

Manufacturing system risk control knowledge matching method based on semantic embedding and clustering analysis

The invention relates to a manufacturing system risk control knowledge matching method based on semantic embedding and clustering analysis, and the method comprises the following steps: collecting and preprocessing risk control text data: collecting unstructured text data of a manufacturing system history record, and obtaining preprocessed risk control text data, constructing a professional corpus for a discrete manufacturing scene; text semantic embedding generation; semantic clustering modeling: performing unsupervised clustering modeling on all semantic vectors, mining semantic association and potential structures between texts, obtaining semantic representations of risk control knowledge through a clustering algorithm, and assisting in generating clustering tags; and a risk knowledge matching mechanism.
Owner:TIANJIN UNIV

Infrared image single target tracking method based on hyperbolic-Euclidean space feature modeling

The invention relates to an infrared image single target tracking method based on hyperbolic-Euclidean space feature modeling, and belongs to the field of computer vision. The method comprises the following steps: firstly, through a parallel VisionMama network and a hyperbolic Transform network, obtaining a space-time visual feature and a hierarchical structure feature of a search frame; secondly, constructing a hierarchical tree based on a Gaussian model, dynamically decoding the hierarchical structure features of the search frame by adopting an attention mechanism, and optimizing a semantic clustering process through KL regularization constraint to obtain enhanced hierarchical structure features; then, fusing the space-time visual features of the search frames and the enhanced hierarchical structure features by using a weighted fusion technology; and finally, reconstructing the fused features, inputting the reconstructed features into a full convolutional network, and positioning a target by adopting weighted focus loss and multi-target regression loss. According to the method, reliable hierarchical feature extraction can be realized, the adaptability of the model to target deformation, shielding and rapid movement is enhanced, and a high-precision and high-robustness solution is provided for infrared single target tracking in an extreme environment.
Owner:KUNMING UNIV OF SCI & TECH

Data marking method and device, equipment and medium

The invention provides a data marking method. The method can be applied to the technical fields of big data and artificial intelligence. The method comprises the steps of obtaining multiple pieces of multi-modal data, preprocessing the multiple pieces of multi-modal data, and generating multiple pieces of preprocessed text data; and performing vector conversion on the plurality of pieces of preprocessed text data to generate a plurality of pieces of feature vector data. Density clustering is carried out on the multiple pieces of feature vector data, a data sets are generated, and each data set comprises a first data label. Semantic clustering is conducted on the a first data labels through semantic analysis, and b second data labels are generated. And presetting a business knowledge graph, and performing knowledge fusion on the business knowledge graph and the b second data tags to generate b target data tags for data marking. The invention further provides a data marking device and equipment, a storage medium and a program product.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

File content semantic clustering method based on graph neural network

The invention discloses an archive content semantic clustering method based on a graph neural network, and the method comprises the following steps: S1, carrying out the word segmentation, denoising and vector expression of an archive text, and generating a text feature vector set; s2, constructing a semantic graph structure model according to the semantic similarity and the reference relationship between the texts; s3, generating a division result with a balanced structure on the semantic graph by adopting a graph division algorithm; s4, performing double-layer node merging on each graph division cluster, and generating a graph structure coarsening result and a mapping relation; s5, inputting the original graph and the coarsened graph into the graph neural network model, and calculating and fusing each layer of semantic representation; s6, cross-layer consistency constraint optimization node semantic representation is introduced, and a unified embedded vector set is generated; and S7, inputting the embedded vector into the clustering model, and outputting a semantic clustering category of the archive text. According to the invention, semantic recognition and automatic grouping of archive contents are realized.
Owner:THREE GORGES HI TECH INFORMATION TECH CO LTD

Dynamic context window customer service quality inspection method and system based on large model

The invention discloses a dynamic context window customer service quality inspection method based on a large model, and the method comprises the following steps: S1, generating a dynamic window: scoring the dialogue round of a customer service and a user, dynamically updating and maintaining a context window according to a scoring result, and carrying out the context reconstruction of a dialogue in the window; s2, multi-dimensional quality inspection: inputting the dialogue text in the dynamic window into a large model, and outputting a preset field through a customized prompt trigger model; s3, clustering attribution: carrying out semantic clustering on related contents output by the large model by adopting a kmeans algorithm and BERT vectorization, and then generating a general description and an operable suggestion for each clustering result through prompt; and S4, result output and application: generating a structured json result containing a multi-dimensional quality inspection result and a clustering result, wherein the structured json result is used for api calling or visual platform display. The method has the advantages that efficient, accurate and multi-dimensional customer service quality inspection can be achieved, and the service quality can be improved and overcome.
Owner:SHENZHEN SKIEER INFORMATION TECH CO LTD

Large-scale knowledge graph visualization method and system

The invention provides a large-scale knowledge graph visualization method and system, and relates to the field of knowledge graph visualization. The method comprises the following steps: acquiring knowledge graph data, and clustering the knowledge graph data through a modularity-based discovery algorithm; obtaining each sub-graph corresponding to the clustered knowledge graph data, and for each sub-graph, selecting a representative node based on a PageRank algorithm or a Leader Rank algorithm; carrying out force-oriented layout on the clustered knowledge graph data through a tree diagram space filling technology; and for the knowledge graph data subjected to the force-oriented layout, distributing priorities and use times of a Barnes-Hut algorithm and a random vertex sampling algorithm according to a preset mode, and dynamically displaying a visualization result in a layered manner through an affine transformation technology. According to the method and the device, the problem that the structural expression clarity of the drawn knowledge graph is greatly reduced due to the fact that semantic clusters and hierarchical organizations in the graph are difficult to accurately present in a traditional visualization method is solved.
Owner:WUHAN UNIV OF TECH +1

Precise matching and distributing method and system for garment styles

The invention discloses a clothing style accurate matching distribution method and system, and relates to the technical field of clothing personalized recommendation, and the method comprises the steps: collecting a clothing matching coupling data set, carrying out the spatial-temporal feature decoupling, forming a style selection feature matrix, carrying out the semantic clustering analysis and manifold space mapping of the style selection feature matrix, and outputting a clothing style selection label. Inputting the clothes style selection labels and the user behavior data into a collaborative filtering recommendation engine, executing label similarity calculation and user behavior modeling, generating a clothes matching candidate set, and performing behavior frequency weighting and attribute preference coupling on the clothes matching candidate set to obtain behavior-preference coupling weight parameters. According to the invention, the clothing style selection label is generated through manifold mapping, and the personalized ability of clothing matching and distribution is improved. And meanwhile, through a space-time cooperation conversion rate estimation model and an improved PageRank algorithm, the matching accuracy of the recommendation result and the actual demand of the user is improved, and full-link accurate matching distribution of the clothing is realized.
Owner:QINSILK COM

Data synchronization method and system for intelligent handheld terminal

The invention discloses a data synchronization method and system for an intelligent handheld terminal, and relates to the technical field of data transmission, and the method comprises the steps: building the connection between a terminal and a server, authenticating the identity of a user and the identity of the terminal, obtaining server data after the identity authentication is completed, comparing the server data with terminal data, extracting difference data, and determining synchronous data; a synchronized data index is generated based on the synchronized data. According to the method, the context modeling capability of the Transform model, the learnable semantic clustering mechanism of the ClusterFormer and the relative position embedding technology in the Method 4 are fused, structuring, clustering perception and semantic enhancement modeling are performed on the synchronous data, intelligent sorting and dynamic priority evaluation of the synchronous data are realized, and the synchronization robustness, efficiency and accuracy are remarkably enhanced.
Owner:SHENZHEN BLOVEDREAM TECH CO LTD

Content auditing abnormity monitoring and early warning method and system based on intelligent alarm suppression

The invention discloses a content auditing abnormity monitoring and early warning method and system based on intelligent alarm suppression, and the method comprises the following steps: obtaining real-time content auditing data collected in multiple dimensions, determining multi-dimensional data, carrying out the data cleaning and standardization processing of the multi-dimensional data, and carrying out the monitoring and early warning of the content auditing abnormity. Storing the multi-dimensional data subjected to data cleaning and standardization processing into a database; and constructing a deep reinforcement learning model, determining data input of the deep reinforcement learning model based on the standardized multi-dimensional data, and driving the deep reinforcement learning model to optimize an alarm mode in a training process. According to the method, hidden violation and semantic evolution trends are recognized through semantic clustering, and the traditional detection bottleneck based on a numerical threshold value is broken through; according to the invention, by realizing cross-cycle trend evolution analysis, early signals before abnormal outbreak can be identified in advance; the method has the cross-event semantic comparison capability, and the false alarm rate and the redundant data volume are remarkably reduced.
Owner:CHONGQING KAIYUAN GONGCHUANG TECH CO LTD

Front-end cache management method, system and equipment for conversation state of lightweight large model and medium

The invention discloses a front-end cache management method, system and device for a lightweight large-model dialogue state and a medium, belongs to the technical field of front-end cache management of a large-model dialogue system, and aims at solving the technical problem of how to overcome the defects that in a traditional scheme, long context cache is low in efficiency, storage redundancy and insufficient in dynamic semantic adaptation capacity, and the large-model dialogue state cannot be managed easily. In order to realize dialogue context volume compression, improve semantic similar request hit rate and reduce cross-end synchronization delay, the adopted technical scheme is as follows: data acquisition and preprocessing: capturing user interaction behaviors in real time through front-end burying points, and performing preprocessing operation on the acquired user behavior data; semantic normalization processing: performing embedded vector conversion and semantic clustering on the text input by the user to generate a unique semantic identifier and a context vector; querying and updating the multi-level cache; and dynamic collaborative updating: dynamically adjusting the cache based on the cache hit rate, the response delay and the user feedback, and optimizing the cache effect in real time.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

QA-driven large-model hierarchical knowledge graph parameterization construction method

According to the QA-driven large-model hierarchical knowledge graph parameterization construction method provided by the invention, the unstructured document is segmented, and the structured question and answer pairs are generated in combination with the language model, so that automatic extraction from the original text to the knowledge unit is realized, and the manual participation cost is reduced; then, semantic affiliation and hierarchical relations are established between the extracted question content and entity categories, atomic nodes and intermediate nodes, systematic hierarchical organization of the knowledge graph is achieved, the structure is clear, and logic is complete. Moreover, a semantic clustering algorithm is adopted to classify bottom nodes, and intermediate nodes are automatically constructed on the premise of meeting graph structure constraint conditions, so that redundancy and repetition are avoided; and finally, through information granularity decomposition driven by question and answer pairs and in combination with atomic node construction, the fine degree of knowledge extraction and the expressive power of graph representation are remarkably improved.
Owner:GUANGDONG TECSUN SCIENCE & TECHNOLOGY CO LTD

Inland ship abnormal behavior sensing method and device based on incremental graph convolutional network

The invention provides an inland ship abnormal behavior sensing method and device fusing multi-modal data based on an incremental graph convolutional network. Comprising the following steps: acquiring laser point cloud data, video image data, AIS data and navigable water level data of an inland ship; the obtained four types of data are preprocessed; video and laser point cloud data fusion: based on a two-stage front and back fusion algorithm, fusing the preprocessed data to obtain a ship image target with accurate three-dimensional information; video and AIS data fusion: fusing the preprocessed video image data and AIS data to obtain a ship target containing ship attribute data; based on the ship feature data, constructing a semantic clustering incremental graph model S-GCN; constructing an incremental graph model M-GCN of multi-modal information association fusion; constructing an incremental graph convolution framework graph based on multi-modal information fusion abnormal behavior perception of an incremental graph model; and matching the feature data to obtain fused multi-modal feature data, inputting the fused multi-modal feature data into a detection framework network, and showing relatively high detection performance, detection accuracy and effectiveness in ship yaw early warning detection, ship bridge crossing early warning detection and ship collision early warning detection tasks.
Owner:TIANJIN RES INST FOR WATER TRANSPORT ENG M O T

Network protocol fuzz testing method based on large language model

The invention discloses a network protocol fuzzy test method based on a large language model, which is characterized in that key information is automatically identified and extracted from a network protocol document by utilizing an advanced natural language processing technology, then a knowledge graph is constructed, and a fuzzy test driver is generated through cue word engineering on the basis of the graph. The method comprises the following steps: firstly, creating a network protocol knowledge graph, processing a related network protocol specification document by using a large model, creating references to all message type entities and protocol related states in the document, and then generating the network protocol knowledge graph by using the information through the large model; then, using the knowledge graph to create clusters from bottom to top, hierarchically organizing data into semantic clusters, and summarizing semantic concepts and topics in advance, so that comprehensive understanding of a network protocol is facilitated; the test driver code is then generated based on the prompt enhancement at the time of the query.
Owner:BEIHANG UNIV

Cross-file information summarization and new knowledge automatic summarization method and system

The invention discloses a cross-file information summarization and new knowledge automatic summarization method and system, and belongs to the technical field of file data processing, and the method specifically comprises the steps: collecting document file data, vectorizing the document file data based on a semantic embedding model, carrying out the reconstruction and semantic clustering of the semantic vector of the document file data, and carrying out the automatic summarization of the new knowledge. The method comprises the following steps: constructing a semantic structure chart which represents a logical relationship between document file data, performing content completion on nodes which are not completely associated in the semantic structure chart, extracting structured knowledge units from the completed semantic structure chart, and generating a new knowledge text based on the structured knowledge units; according to the method, the limitation of low efficiency of knowledge splitting and manual summarization between traditional document files is solved, the automation degree of knowledge discovery is remarkably improved, and the method is suitable for scenes of knowledge base construction, domain rule extraction, domain knowledge discovery and the like.
Owner:GUIZHOU BLUE DREAM FACTORY TECH CO LTD

Semantic clustering for unlimited context window sizes for sequence processing models

Systems and methods are provided for semantic clustering for arbitrarily long context windows for machine-learned sequence processing models. A computing system can obtain a context sequence. The computing system can determine a plurality of subsequences of the context sequence. The computing system can determine, using a machine-learned semantic embedding model, a semantic embedding for each subsequence. The computing system can determine, based on the semantic embedding, a plurality of semantic clusters. The computing system can generate, using a machine-learned sequence generation model and based at least in part on the semantic clusters, an output sequence.
Owner:GOOGLE LLC

Cross-modal retrieval method based on comparative learning and balanced hash coding

The invention discloses a cross-modal retrieval method based on comparative learning and balanced Hash coding, which comprises the following steps: respectively extracting modal specific features and modal shared features from input multi-modal data, aligning the modal specific features of different modals through comparative learning, modal specific features and modal sharing features are optimized by using quantization loss based on optimal transmission; performing binarization operation on the modal specific feature and the modal shared feature to generate a modal specific hash code and a modal shared hash code; semantic clustering is carried out through a K-means algorithm based on the modal shared hash code to generate a semantic index; candidate samples are coarsely screened in the cross-modal retrieval stage through semantic indexes, then fine-grained comparison is conducted through generated modal specific hash codes and modal shared hash codes, and efficient cross-modal retrieval is achieved. According to the method, semantic alignment of different modes is realized by utilizing comparative learning, and the distribution characteristics of hash codes are optimized through the quantization loss based on optimal transmission, so that the retrieval performance is improved.
Owner:SOUTH CHINA UNIV OF TECH

Neural network compression system based on lexical attention score dynamic pruning

A neural network compression system based on lexical attention score dynamic pruning comprises an input module, a mask generation module and a dynamic reasoning module, attention scores are directly introduced into a pruning decision to quantify the importance of an internal structure of a model, fine pruning based on real attention distribution is achieved, and the accuracy of pruning is improved. The distortion problem of a traditional weight amplitude-based method is avoided; according to the method, lexical elements are divided through semantic clustering, independent masks are generated in each class, and class-level pruning granularity is constructed, so that a model structure is more adaptive to input semantic features, performance stability is kept under a high pruning rate, lexical element classes are identified based on a KNN algorithm, and the masks are dynamically called; sparse strategy switching during reasoning is realized, reasoning efficiency and model precision are both considered, and a new dynamic control path is provided for lightweight reasoning of a large model.
Owner:SHANGHAI JIAOTONG UNIV

Process contrastive analysis system and method based on standard semantic analysis and computer readable recording medium

The invention relates to a technology comparative analysis system and method based on standard semantic analysis and a computer readable recording medium. According to the system, dimension reduction processing is carried out on process data and standard texts through a semantic clustering module, and a semantic mapping relation between processes and standard terms is constructed; generating a contrast set by using a process operation standard threshold construction module and a convolutional network; and calculating a compliance score by adopting a contrast learning module. The system supports dynamic updating of a multi-dimensional standard knowledge graph, analyzes texts and recognizes semantic equivalent pairs through a natural language processing technology, achieves intelligent comparison and deviation recognition of process data and standard rules, automatically generates compliance reports and early warning information, and finally achieves digital management of enterprise standard full life cycles. And automation and accuracy of compliance analysis are improved.
Owner:HIGH QUALITY STANDARDIZATION RES INST (SHANDONG) CO LTD