Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

24 results about "Semantic clustering" patented technology

Semantic clustering helps your company discover gaps in your content to enrich your customer’s experience. Inbenta’s Semantic Clustering groups semantically equivalent search queries — words, phrases and sentences — into clusters based on meaning. The higher the number of questions, words and phrases with a similar meaning, the greater the cluster.

Semantic clustering of messages

Example systems, methods, and computer-readable media are disclosed. In an example method, a first outbound text message is transmitted via a message broker of a messaging platform from a client to a plurality of recipients. In response to the first outbound message, a plurality of inbound text messages is received, via the message broker, from the plurality of recipients. A first grouping of the plurality of inbound text messages is determined, the first grouping associated with one or more recipients of the plurality of recipients. The first grouping is presented to the client. A second outbound text message is transmitted, via the message broker, from the client to the one or more recipients of the plurality of recipients. The second outbound text message is generated based on the first grouping. The message broker is in communication with a first messaging service and a second messaging service different from the first messaging service. The first outbound text message is transmitted via the first messaging service. A first inbound text message of the plurality of inbound text messages is received via the second messaging service. Each inbound text message of the plurality of inbound text message is addressed to a long-code telephone number generated by the messaging platform and uniquely associated with the client by the messaging platform.
Owner:COMMUNITY COM INC

A privacy data desensitization method for a scientific and technological achievement transformation platform

ActiveCN121659360BFeature vectorSemantic clustering
The application relates to the technical field of data security, in particular to a privacy data desensitization method for a scientific and technological achievement transformation platform, which comprises the following steps: collecting scientific and technological achievement data from the scientific and technological achievement transformation platform, and performing word segmentation and keyword extraction on the text content; performing semantic clustering on the keywords based on the feature vectors of the keywords to construct a generalization hierarchical tree; weighting and fusing the sensitivity and rarity of the keywords to obtain data privacy; dividing equal-length intervals according to the index size of the data privacy, and assigning an adaptive K value to each group for generalization by using a K-anonymity algorithm to perform generalization desensitization processing on the keywords in the scientific and technological achievement data. The application aims to ensure the privacy of high-risk sensitive data in the scientific and technological achievement data and avoid excessive generalization and loss of data effectiveness.
Owner:BEIJING INFOSOFT CO LTD

Multilingual generative retrieval method based on cross-language semantic compression

ActiveCN120892582BData setDocument Identifier
This invention relates to a multilingual generative retrieval method based on cross-language semantic compression, belonging to the field of information retrieval technology. The invention includes the following steps: constructing a multilingual document retrieval dataset; extracting keywords from multilingual documents from multiple perspectives using a keyword extraction model, and calculating the extracted keywords using semantic similarity to construct a similarity matrix; performing semantic clustering based on the similarity matrix, representing clusters using atomic IDs, and then assigning document identifiers to each multilingual document by the cluster containing the keywords; in the inference stage, after inputting a query, employing a dynamic multi-complement constraint decoding method, gradually narrowing the decoding range of the document identifier in the current step based on the decoding results of previous steps, thereby obtaining the final document identifier. The retrieval capability of this invention is significantly improved compared to other models.
Owner:KUNMING UNIV OF SCI & TECH

A large model training data deduplication method based on semantic clustering

PendingCN122153263ASemantic analysisBiological modelsSemantic vectorSemantic clustering
The application discloses a kind of big model training data deduplication method based on semantic clustering.The method is first filtered by Hash matching and MinHash structure to completely repeated and locally repeated text;Subsequently, a pre-training semantic encoding model is used to generate a deep semantic vector, and a density clustering method is used to construct a semantic cluster, converting global high-complexity comparison into local retrieval within the cluster;Then, combined with the efficient neighbor search based on FAISS, the semantic similar samples within the cluster are grouped and deduplicated;Finally, an information entropy sensing mechanism is introduced, and representative samples are selected according to the content complexity of the samples, to significantly reduce redundancy while maintaining corpus diversity.The application can effectively reduce the risk of model memory caused by data duplication, improve training efficiency, and reduce data processing cost, suitable for various large-scale corpus construction and privacy-sensitive scenarios.
Owner:ZHEJIANG UNIV +1

An advertisement putting rhythm control method based on crowd behavior prediction

PendingCN122288794AClustered dataSemantic clustering
A method for controlling the pace of advertising delivery based on group behavior prediction, relating to the field of advertising control, is proposed. It involves generating topic cluster data through semantic clustering of cross-platform public opinion data, and generating cross-platform topic alignment data; identifying bridging proxy data that acts as a propagation carrier between platforms based on the cross-platform topic alignment data, and calculating bridging strength data; configuring the bridging strength data into a preset propagation dynamics model to generate diffusion prediction results data; combining delivery log data and cross-platform topic alignment data to generate delivery coupling degree data; and generating delivery pace control plan data based on the delivery coupling degree data when the negative public opinion spillover risk data is greater than or equal to a first preset threshold and the predicted arrival time data falls within a preset monitoring window period. This improves the stability and accuracy of advertising delivery control.
Owner:SHANGHAI INTERNATIONAL STUDIES UNIVERSITY

Agent-driven semantic clustering directed fuzzing method and apparatus

The embodiment of the application discloses an agent-driven semantic clustering directed fuzzing method and device, and relates to the technical field of software security testing.The method specifically comprises the following steps: firstly, a target point is acquired and is merged according to functions; secondly, the control flow graph of each function is analyzed before testing, key predicates are extracted, and a constraint signature is generated for the target point; the constraint signature is divided into semantic clusters based on a logical relationship, and combined constraint signatures of the semantic clusters are formed in combination with calling contexts; then, in the testing runtime, the semantic clusters are taken as units, seeds and energy are dispatched according to the combined constraint signatures of the semantic clusters, the state is updated according to feedback, and it is determined whether the semantic clusters are stagnant; when the semantic clusters are stagnant, the input is directed to be adjusted to generate a new seed and to be executed based on the combined constraint signatures of the semantic clusters and the path of the current seed; finally, the state of the semantic clusters and the historical strategy are updated according to the execution result of the new seed, and the above-mentioned dispatching and generating process is iteratively executed, so that efficient and accurate vulnerability testing is realized.
Owner:XIAMEN UNIV OF TECH

A distributed file system based on content semantic hash

PendingCN122412378ADistributed File SystemFile system
The application relates to a distributed file system based on content semantic hash and belongs to the technical field of data storage. The system breaks through the retrieval and storage limitation of traditional file systems depending on file names or file paths by means of a content semantic hash technology, realizes semantic-driven management of data in combination with a distributed cluster architecture, and solves problems such as high redundancy of heterogeneous data, low retrieval precision, and low efficiency of storage resource scheduling. The system takes a hash index layer as a core support, extracts deep semantic features of files through a semantic processing layer, maps the features to generate unique semantic hash values by the hash index layer, and realizes intelligent deduplication, semantic clustering storage and accurate retrieval on the basis.
Owner:BEIJING INST OF COMP TECH & APPL

Sentence semantic clustering compression method based on heterogeneous kv cache, electronic device, program product

ActiveCN121980031BSemantic representationContextual reasoning
The application discloses a sentence semantic clustering compression method based on a heterogeneous KV cache, electronic equipment and a program product. The method comprises the following steps: reserving the first T tokens of a given query on a GPU, and splitting the remaining tokens to obtain S sentences; for each sentence, taking the mean of the Key vector corresponding to the token as the sentence center, calculating the similarity between the Key vector corresponding to each token and the sentence center, calculating the GSA weight of each token according to the similarity, and calculating the semantic representation according to the GSA weight; and clustering the semantic representations of the S sentences to obtain C cluster representations. The application improves the accuracy and efficiency of long context reasoning, and reduces the long sequence reasoning delay and memory pressure under the limited KV budget.
Owner:UESTC (SHENZHEN) ADVANCED RES INST

An artificial intelligence-based semantic clustering conversation intelligent scheduling system and method

The application discloses an intelligent scheduling system and method for semantic clustering conversation based on artificial intelligence, relates to the technical field of semantic analysis, and synchronously acquires voice call data and online customer service text flow, adopts an attention mechanism to fuse a joint semantic vector and a user value mapping vector, and generates a context coding vector in combination with a decay factor; an initial topic node is constructed based on the context coding vector, a node weight is dynamically updated through a conversation frequency decay factor, and a topic correlation graph with a weighted edge is generated; for a historical conversation associated with a graph node, a three-dimensional performance model is constructed, a performance score of a customer service reply technique is analyzed, and a knowledge base sorting update is triggered; a value index change rate of a topic node is monitored in real time, when a sudden increase of a high net worth user group associated conversation is detected, a fuse mechanism is automatically triggered and a dedicated agent is dynamically allocated. The application improves the customer service response quality and improves the service experience of users.
Owner:WESHINE CO LTD

A large model-based text generation method, electronic equipment and storage medium

The application relates to the technical field of large model retrieval, in particular to a text generation method based on a large model, an electronic device and a storage medium, the method comprises the following steps: rewriting an initial question input by a user, performing sorting processing on a plurality of relevant documents obtained based on the rewritten question, screening out target documents, performing key information extraction on each target document, performing semantic clustering processing on the extracted key information segments based on a dynamic clustering quantity, obtaining a plurality of clustering clusters, obtaining cluster representative contents corresponding to each clustering cluster after deduplication processing, to generate pretreatment texts, matching a target report template corresponding to the initial question from a preset report template library, and generating a target report text in a format corresponding to the target report template according to the pretreatment texts; the application realizes intelligent generation of a whole process from fuzzy input of a user to a precise and structured report, and significantly improves the accuracy, relevance and format standardization of report contents.
Owner:BEIJING YUCHEN SHIMEI SCI & TECH

An artificial intelligence english corpus quick matching system for teaching training purposes

The application discloses an artificial intelligence English corpus rapid matching system for teaching training, which comprises, sequentially connected, a corpus preprocessing and multi-view index construction module, an input understanding and intention recognition module, a mixed recall and rearrangement module, a semantic clustering and knowledge unit generation module, and a correlation degree calculation and decomposition display module; the correlation degree calculation and decomposition display module is connected with an exercise generation and evaluation module and a learning portrait and recommendation module; and the corpus preprocessing and multi-view index construction module and the semantic clustering and knowledge unit generation module are connected with a subword alignment and word family aggregation module. The application realizes personalized adaptation based on a learning portrait, can provide targeted corpus, knowledge units and exercise recommendations according to the learning level, weak knowledge points and other characteristics of different users, adapts to the needs of different learning groups, and improves the efficiency and quality of English teaching training.
Owner:GUANGDONG UNIV OF TECH +1

A data de-sensitization system and method

PendingCN122333517AEngineeringMulti source data
This invention discloses a data anonymization system. A data classification module categorizes multi-source data in a database, obtaining structured and unstructured multi-source data. A sensitive data identification module identifies structured sensitive information from the structured multi-source data and uses semantic clustering to identify unstructured sensitive information from the unstructured multi-source data. The unstructured sensitive information is then converted into structured sensitive information through basic parsing. A data privacy module classifies the structured sensitive information according to preset information security levels, obtaining structured sensitive information at each security level. A matching data anonymization method is then used to anonymize the structured sensitive information at each security level, resulting in anonymized multi-source data. This invention achieves one-stop operation for sensitive data management, balancing data security and availability, meeting relevant compliance requirements, effectively reducing the risk of data leakage, and is suitable for sensitive data protection in multiple business scenarios.
Owner:NAVAL UNIV OF ENG PLA

Intelligent research and judgment method and system for false alarm of multi-source static analysis alarm

PendingCN122364046AData setLinguistic model
This application belongs to the field of software security and program analysis technology, specifically disclosing a method and system for intelligent false alarm judgment of multi-source static analysis alarms. Through this application, raw warning data is mapped to a standardized alarm representation with a unified field structure; a target alarm data set is generated based on the deduplicated raw warning data; semantic representation learning is performed on code fragments; semantically similar alarm data are clustered; composite prompt information is constructed based on candidate knowledge entries and target code fragments corresponding to representative alarm data; and intelligent analysis and judgment are performed based on a large language model. Through the above methods, a series of processes are performed on the multi-source raw warning data generated by static analysis, including unified representation, explicit deduplication, semantic clustering, knowledge retrieval, false alarm analysis, and judgment. This achieves an integrated processing flow from alarm input to result distribution, without relying on project historical information, thereby effectively improving the efficiency and accuracy of false alarm analysis and judgment.
Owner:HUAZHONG UNIV OF SCI & TECH

A method, apparatus, device and medium for text clustering

ActiveCN116578702BImplement automatic clusteringEfficient semantic clusteringDigital data information retrievalNatural language data processingAlgorithmWord list
The application provides a text clustering method, device and equipment and readable medium, the method comprises the following steps: establishing a vocabulary and calculating the word vector of each word in the vocabulary; obtaining the text vector of each text to be clustered and forming a text vector set, and calculating the distance between each two text vectors in the text vector set; randomly selecting a threshold number of text vectors in the text vector set as candidate center vectors, and dividing the text vectors into two categories by taking each two text vectors as a group and sequentially taking the candidate center vectors as center vectors; selecting the center vector with the maximum confusion degree in the center vector group with the minimum confusion degree in each division and the corresponding classified text vector, and repeating the previous step with the selected text vector until a preset condition is reached. By using the scheme of the application, efficient semantic clustering of short texts can be achieved, and automatic clustering of short texts can be achieved while fully preserving the semantic and sequence information of the texts.
Owner:JINAN INSPUR DATA TECH CO LTD

Contract information extraction method, device and equipment based on multi-level prompt word instruction

This invention discloses a method, apparatus, and device for extracting contract information based on multi-level prompt word instructions. The method includes: acquiring the digital signal of the contract document to be processed; identifying business type tags through layout analysis and semantic clustering; then, using an instruction scheduling engine, calling and reconstructing multi-level instruction modules from a preset instruction library based on the tags to assemble a complete multi-level instruction stream; finally, injecting the complete multi-level instruction stream into an intelligent processing module, driving the intelligent processing module to complete the contract information extraction according to the instruction stream constraints, and outputting standardized structured data after verification and auditing. The multi-level instruction module is a three-layer architecture consisting of a decoupled system meta-instruction layer, an atomic rule configuration layer, and a semantic reasoning audit layer. These layers are independent of each other, and lower-level modules cannot modify the upper-level constraint rules. This invention solves the problems of instruction bloat, subject ambiguity, black-box unauditability, and weak generalization ability.
Owner:SHANGYANG TECH CO LTD

A large language model key-value cache compression method and system based on semantic cluster center offset

PendingCN122347181AVideo memoryLinguistic model
This invention relates to a key-value caching compression method and system for large language models based on semantic cluster center offset, belonging to the field of large language model inference acceleration technology. The method includes: a pre-filling stage, where semantic clustering is performed on the key matrix and cluster center vectors are extracted; the offset between the key vectors and the cluster centers is calculated; unstructured channel pruning and low-bit quantization are performed on the offsets; compressed data is stored in off-chip video memory; and cluster centers are stored in on-chip shared memory. In the decoding stage, compressed data is read from off-chip video memory; the offsets are dequantized to recover the data; approximate key vectors are reconstructed using the on-chip cluster centers; and attention calculation is performed. This invention maintains semantic integrity under high sparsity by transforming the compressed object from the original vector into an offset from the semantic cluster centers, eliminating the truncation and destruction of low-bit quantization by large numerical outliers, and achieving lossless superposition of pruning and quantization.
Owner:CHONGQING UNIV

Metal mine dispatching large model fine-tuning data screening method and device and storage medium

PendingCN122432657AFeature vectorData set
The application relates to the technical field of artificial intelligence and intelligent mines, and provides a metal mine scheduling large model fine-tuning data screening method, a device and a storage medium, which comprises the following steps: based on the normalized quality score, the mutual information correlation degree score and the high-dimensional semantic consistency rearrangement evaluation score of a to-be-screened fine-tuning data set, and obtaining the comprehensive evaluation score of the to-be-screened fine-tuning data set; taking the comprehensive evaluation score as a quality evaluation dimension to obtain a candidate fine-tuning data set containing a multi-dimensional score label; obtaining a graph-level high-dimensional semantic topology feature vector of the candidate fine-tuning data set; performing dimension reduction processing on the graph-level high-dimensional semantic topology feature vector to obtain a dimension-reduced semantic feature vector; based on the semantic clustering result of the semantic feature vector, determining a plurality of scheduling fine-tuning data sets from the candidate fine-tuning data set, and generating a target quality scheduling fine-tuning data set with balanced structure distribution and covering multiple scheduling scenarios based on the plurality of scheduling fine-tuning data sets. In this way, the data quality of the fine-tuning data set is improved.
Owner:CENT SOUTH UNIV

Method for adaptive extension of building element dictionary based on recognition feedback and semantic clustering

PendingCN122368551AFeature vectorEngineering
This invention relates to the field of data processing technology and discloses an adaptive expansion method for a building element dictionary based on recognition feedback and semantic clustering. The method includes: acquiring image samples of building elements to be identified and the current version of the building element dictionary, and extracting feature vectors; calculating the matching score between the building element image samples and each element, and determining whether a stable match or no match is achieved; collecting samples determined to be unmatched into an unmatched pool and performing quality filtering; performing density clustering analysis on the image feature vectors of the filtered samples to obtain at least one candidate cluster; automatically determining the operation to be performed on the candidate clusters through a decision function; generating new unique element identifiers for candidate clusters determined to be generated as new elements, and updating the building element dictionary version. This invention achieves traceability and adaptive expansion of the building element dictionary; simultaneously controlling the retrieval consistency and version rollback after dictionary updates, avoiding the spread of errors caused by direct automatic database entry.
Owner:SICHUAN PROVINCIAL ARCHITECTURAL DESIGN & RES INST

A Key-Value Caching Compression Method Based on Importance Awareness, Dynamic Hierarchy, and Clustering Fusion

This invention belongs to the field of large language model inference optimization, and discloses a key-value (KV) cache compression method based on importance-aware dynamic hierarchical and clustering fusion. By combining importance-aware eviction and global semantic clustering fusion mechanisms, it addresses the problem of excessive GPU memory overhead in long text inference for autoregressive large language models based on the Transformer architecture, achieving high-fidelity long text inference under limited GPU memory conditions. The global semantic clustering algorithm captures cross-paragraph semantic relationships in long contexts, achieving redundant background compression and core semantic preservation, overcoming the obstacles of existing merging methods' difficulty in achieving global semantic awareness and static strategies' inability to adapt to dynamic changes. From the perspective of sliding window incremental merging, newly generated tokens are temporarily stored and similarity retrieval is updated, maintaining the local contextual coherence in the decoding stage, avoiding information loss due to premature compression, and improving the overall performance of long text inference while balancing storage efficiency and semantic fidelity.
Owner:DALIAN UNIV OF TECH

Multimodal network vulnerability false alarm static analysis method based on large language model

The application discloses a kind of multi-modal network vulnerability false alarm static analysis methods based on large language model, comprising the following steps: feature library construction, based on vulnerability standard data is generated with the feature code fragment of matching vulnerability description using large language model, and after multidimensional score screening, vulnerability feature library is constructed;Index construction, semantic clustering is carried out to feature code fragment and constructs hierarchical retrieval index, and is associated with vulnerability knowledge graph;Detection, the semantic vectorization processing is carried out to the code to be detected, is matched in hierarchical retrieval index based on semantic vector, recall normalized vulnerability knowledge according to matching result, the recalled vulnerability knowledge is input into large language model with the code to be detected and carries out vulnerability detection.The application is organically fused by feature code generation, hierarchical index construction and adaptive retrieval reasoning, solves the problem that vulnerability feature is dispersed, homologous sample is difficult to aggregate, and retrieval link information is insufficiently associated, significantly improves the accuracy and efficiency of vulnerability detection.
Owner:BEIJING UNIV OF POSTS & TELECOMM

An embedding vector acquisition method, apparatus, electronic device, and storage medium

This invention provides an embedding vector acquisition method, apparatus, electronic device, and storage medium, relating to the field of machine learning technology. The method includes: for each object from which an embedding vector is to be generated, acquiring a first training vector obtained by training the embedding vector based on the object's training samples; performing semantic clustering on the obtained first training vectors to obtain at least one vector cluster; for each vector cluster, acquiring a preset vector as the initial embedding vector for that vector cluster; using the training samples of the objects corresponding to each first training vector in the vector cluster, training the initial embedding vector of the vector cluster to obtain a second training vector for the vector cluster; and for each object, acquiring the object's embedding vector based on the second training vector of the vector cluster to which the object's first training vector belongs. This method can improve the accuracy of the obtained object embedding vectors.
Owner:BEIJING QIYI CENTURY SCI & TECH CO LTD

Multi-agent discussion method, apparatus, device, and medium

The application provides a multi-agent discussion method, device, equipment and medium, the method comprises the following steps: generating a decision factor set containing a preset number of nodes according to discussion topic information, performing semantic clustering, hierarchical labeling and causal relationship analysis on each node in the decision factor set, and constructing a directed acyclic influence graph; performing semantic retrieval on the vector database according to the discussion topic information, and performing node retrieval on the directed acyclic influence graph according to the innovation ability of the agent; generating a speech prompt word according to the results of double retrieval, the six-dimensional prompt of the agent, the role setting and the discussion context, and calling a large language model to obtain the speech proposal of each agent; determining the comprehensive utility score of the speech proposal of each agent to determine the speech agent of the current round from each agent and call the speech of the corresponding speech agent. The application overcomes the homogenization defect of the agent role, ensures that the speech content is both creative and consistent with the business logic, and improves the efficiency and quality of group discussion.
Owner:TSINGHUA UNIVERSITY

Intelligent interaction method, device and equipment of cross-modal data and storage medium

The application relates to the technical field of human-computer interaction and digital film and television cross, and provides a cross-modal data intelligent interaction method and device, equipment and a storage medium. According to the modal type of data, the received plot data set is subjected to feature extraction to obtain an initial feature set of each modal, and the initial feature set is projected to a preset common semantic space to generate a cross-modal plot feature sequence aligned in the semantic space; according to a preset semantic community set, the cross-modal plot feature sequence is subjected to clustering analysis and hierarchical compression coding to generate a plot semantic coding sequence; and according to a preset coding element priority, the plot semantic coding sequence is subjected to hierarchical decoding and time domain synchronous fusion rendering to generate an interactive plot data stream. Through modal alignment, semantic clustering, hierarchical coding, priority scheduling and time domain synchronous fusion, the application realizes low-delay and semantic-consistent interactive plot presentation.
Owner:GUANGZHOU HAND IN HAND INTERNET CO LTD