Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

316 results about "Index term" patented technology

An index term, subject term, subject heading, or descriptor, in information retrieval, is a term that captures the essence of the topic of a document. Index terms make up a controlled vocabulary for use in bibliographic records. They are an integral part of bibliographic control, which is the function by which libraries collect, organize and disseminate documents. They are used as keywords to retrieve documents in an information system, for instance, a catalog or a search engine. A popular form of keywords on the web are tags which are directly visible and can be assigned by non-experts. Index terms can consist of a word, phrase, or alphanumerical term. They are created by analyzing the document either manually with subject indexing or automatically with automatic indexing or more sophisticated methods of keyword extraction. Index terms can either come from a controlled vocabulary or be freely assigned.

Knowledge discovery based on indirect inference of association

Techniques for knowledge discovery based on indirect inference of association are presented. A data management component (DMC) can determine and extract, in a structured format, entities, relationships between entities, and concepts relating thereto in documents, based on analysis of information in the documents and / or keywords relating to concepts, to generate an association inference model. Using artificial intelligence techniques, DMC can embed the entities and relationships to a common representation to generate and train a scoring model that can be used to evaluate and score similarity strength between entities, including entities that do not have a known relationship, and can predict or infer relationships, including indirect relationships, between entities or between concepts. In that regard, DMC or user can evaluate concept-level scores to determine a level of relationship between concepts. DMC can feedback information from the scoring model or evaluation to update the association inference model.
Owner:AT&T INTELLECTUAL PROPERTY I L P

A development data security management method and system based on a small program

This invention discloses a method and system for secure management of development data based on mini-programs, relating to the field of electronic data processing technology. The method includes analyzing and processing mini-program development data to be managed to obtain data conforming to a management format. First, the invention removes missing and duplicate values ​​from the mini-program development data. Then, it performs format conversion and classification on the preprocessed data to ensure its orderliness. Next, it performs feature analysis on different types of mini-program development data to determine the correlation between them, obtaining mini-program development data with the same virtual tags. Finally, it analyzes the mini-program development data with the same virtual tags to construct a keyword index. When a certain type of mini-program development data is needed, simply searching for keywords will retrieve the data's storage location, making data retrieval and use more convenient and efficient.
Owner:NINGBO SANSAN CHENGJIU TECHNOLOGY CO LTD

A long text matching method combining noise filtering and divide-and-conquer strategy

ActiveCN117216189BSemantic analysisSpecial data processing applicationsTimed textSentence similarity
This invention discloses a long text matching method combining noise filtering and a divide-and-conquer strategy. The method includes: constructing a long text matching model, which comprises a keyword extraction layer, an association extraction layer, and a filtering layer. The keyword extraction layer extracts keywords from the text; the filtering layer filters noise from the text based on sentence similarity to obtain a denoised text sequence; and the association extraction layer further removes keywords from the denoised text sequence to obtain the remaining associated text. The long text matching model is trained with the optimization objective of minimizing a set overall loss function, which reflects the global matching distribution and combines the keyword and association matching distributions. For the target text, real-time text matching is performed using the trained long text matching model. This invention improves the generalization ability and accuracy of text matching.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

A data bloodline collection method and device based on a Gbase stored procedure

The application discloses a data bloodline collection method and device based on a Gbase stored procedure, and the method comprises the following steps: removing fixed keywords of a stored procedure definition process body; capturing a stored procedure variable through a regular expression and converting the stored procedure variable into a Key-Value form; replacing a variable name with a variable value according to the Key; processing a loop structure and a branch structure, and processing a self-defined label; and performing syntax compatibility and replacing or removing SQL statements according to some keywords. The application processes some general keywords of the stored procedure, removes redundant information irrelevant to the data bloodline, processes some syntaxes specific to the Gbase in a compatible manner, so that the syntaxes can be parsed by open-source SQL tools, and the application is adapted to the stored procedure of the Gbase, so that the application can help some manufacturers using the Gbase stored procedure to process data in an automatic form to the data bloodline, and the application is convenient for data warehouse construction.
Owner:HUNAN AEROSPACE INFORMATION CO LTD

A data processing method, apparatus, device, and medium

This application provides a data processing method, apparatus, device, and medium. The method includes: acquiring a first text containing business text data; performing a risk assessment on the first text to obtain a risk category result corresponding to the first text; if the risk category result is a first risk category, acquiring keywords of the business text data and searching for the keywords in a standard database; if a request text matching the keywords is found in the standard database, determining the feedback text corresponding to the request text as the business processing result corresponding to the business text data; if no request text matching the keywords is found in the standard database, performing text search processing on the business text data in a target knowledge graph to obtain a business processing result matching the business text data. Implementing this application embodiment can improve the security of text data.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Multi-channel third-party site promotion method and system with real-time effect feedback

The present application relates to the technical field of website information analysis, and discloses a multi-channel third-party site promotion method and system with real-time effect feedback, which comprises collecting multi-channel promotion data and preprocessing, quantifying the dynamic contribution weight of each channel in the user conversion path, restoring the cross-channel user behavior time sequence path, jointly optimizing the keyword bidding, creative material, delivery time period and budget allocation, managing the third-party sites according to the conversion efficiency, and realizing the closed-loop optimization of multi-channel cooperation through execution verification, difference analysis and obstacle factor identification. Thus, the promotion efficiency and the return on investment can be improved.
Owner:DINGGE FILM & TELEVISION CULTURE (TIANJIN) CO LTD

Information tracing method and device, electronic equipment and storage medium

This invention relates to the field of information tracing technology, and discloses information tracing methods, devices, electronic devices, and storage media. The invention first disassembles internet information to obtain combined features formed by the combination of event subject characteristics and event description characteristics, along with corresponding time parameters, constructing a dynamically searchable disassembled feature library. After acquiring the information to be traced, based on its event subject, event description, and event occurrence time, the disassembled feature library is retrieved, and a query feature cluster is constructed based on the retrieved combined features. A thorough tracing search is then performed based on the query feature cluster. The earliest publication time of the traceable content is compared with the currently recorded tracing time. If the earliest publication time is earlier than the tracing time, the tracing time is updated to the earliest publication time, and the feature information of the traceable content is added to the information to be traced, for the next iteration. In this way, each iteration supplements earlier features and new keywords, thus effectively tracing back to an earlier, true initial source.
Owner:BEIJING ZHIHUI XINGGUANG INFORMATION TECH CO LTD

A Method and System for Constructing an Industry Knowledge Base Based on Text-Based Data Synthesis

This invention relates to the field of knowledge base construction technology, specifically disclosing a method and system for constructing an industry knowledge base based on text-based data synthesis. The method includes entity recognition of multi-source text data, constructing a knowledge point sequence, and marking key knowledge points in the knowledge point sequence; expanding the key knowledge points based on a large language model to generate an expanded knowledge point set; constructing a knowledge graph based on the expanded knowledge point set, and constructing a preset number of knowledge connection paths in the knowledge graph; classifying the knowledge connection paths and inserting them as knowledge skeletons into the industry knowledge base. When dealing with massive amounts of multi-source text data, this invention expands the keywords based on a large language model to obtain expanded knowledge points, then constructs a knowledge graph, extracts node paths from the knowledge graph, uses the node paths as knowledge skeletons, classifies them, and stores them in the knowledge base. While the resulting knowledge base still contains a large amount of information, it significantly simplifies its size.
Owner:LANYUN NET

Dialogue processing methods, training methods and devices for question rewriting models

This disclosure proposes a dialogue processing method, a training method for a question rewriting model, and a device thereof. The dialogue processing method includes: when information is missing in the user question, inputting the user question and historical dialogue content into the encoding layer of the question rewriting model to obtain the semantic representation vectors of each word in the historical dialogue content and the semantic representation vectors of each word in the user question, as well as the keyword scores of each word in the historical dialogue content; then, the decoding layer accurately selects keywords semantically related to the user question from the historical dialogue content based on the semantic representation vectors of each word in the user question, the semantic representation vectors of each word in the historical dialogue content, and the keyword scores, and explicitly generates a fully rewritten user question based on these keywords. This accurately supplements the missing information in the user question, achieving completeness and readability of the user question content and improving the user experience.
Owner:JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD

Global industry data intelligent integration and analysis method based on multi-source heterogeneous data

PendingCN122451476ASemantic matchingEngineering
The present application relates to the technical field of data processing, and especially relates to a global industry data intelligent integration and analysis method based on multi-source heterogeneous data. The present application collects multi-source heterogeneous data and carries out normalization processing, constructs and trains an industry data analysis model, retrieves global industry data according to a demand keyword of an export commodity through the industry data analysis model to obtain relevant data entries, carries out alignment processing on the original words of the retrieved relevant data entries and the demand keyword, determines the number of data entries that are correctly aligned, calculates the alignment accuracy, determines whether the analysis of the global industry data meets the standard based on the alignment accuracy, adjusts the matching similarity threshold, reduces the lower limit of entity matching confidence, increases the feature matching step, adjusts the keyword filtering intensity and increases the semantic matching depth. The present application effectively realizes accurate monitoring of the integration and analysis of global industry data, and effectively improves the analysis efficiency of global industry data.
Owner:HUASHANG INTERNATIONAL TECHNOLOGY (GUANGZHOU) CO LTD

Intelligent retrieval system for unstructured documents

The application relates to the technical field of document retrieval, in particular to an intelligent retrieval system for unstructured documents. The system comprises a data acquisition module for acquiring unstructured documents; a document feature analysis module for determining representative feature values in combination with keyword semantic importance, paragraph quantity and frequency, and constructing a theme consistency feature vector based on local and global dimensional theme distribution; a document classification module for clustering by comprehensively calculating and measuring distance of themes, keywords and consistency features, and selecting representative documents to construct a knowledge graph; and a retrieval module for generating a retrieval result based on the knowledge graph in combination with a large language model. The application solves the problem of serious homogenization of unstructured document retrieval results, improves the efficiency of intelligent retrieval of unstructured documents by clustering and deduplication and combining with a knowledge graph to enhance semantic association.

A complaint root cause tracing method and device

ActiveCN115757833BPattern matchingData mining
This application provides a method and apparatus for tracing the root causes of complaints. The method includes: acquiring complaint content data; processing the complaint content data using a preset causal extraction pattern matching model to obtain complaint keywords; extracting keywords from the complaint keywords to obtain complaint cause keywords; constructing a complaint graph based on the complaint keywords and complaint cause keywords; and tracing the root causes of complaints based on the complaint graph to obtain the root cause tracing results. It is evident that implementing this method can trace the root causes of complaints, identify the root problems that led to the complaints, and thus fundamentally resolve such complaint issues. This is beneficial for improving the efficiency and effectiveness of complaint management and effectively enhancing the consumer experience.
Owner:PING AN BANK CO LTD

Search system

ActiveUS12645307B2Unstructured textual data retrievalSpecial data processing applicationsMedicineUser input
A system for searching contents within a database of textual contents is described. Said contents and the corresponding databases may be of any kind such as movie titles, music titles, song titles, scientific terms, medical titles / terms, titles formed of one or more sequences of symbols, a list of telephone numbers, a contact list, etc. Upon providing one or more keyword, the search system provides a list one or more of corresponding contents to a user. According to one aspect, after selecting a presented content, by the user, a process corresponding to selected content is executed by a processor.
Owner:GHASSABIAN BENJAMIN FIROOZ

A keyword matching method, an electronic device cluster, and a program product

The application relates to the computer technical field, and provides a keyword matching method, an electronic device cluster and a program product, and further provides a computer readable storage medium. The keyword matching method provided by the application is applied to an electronic device, and the method comprises the following steps: extracting all continuous words satisfying the length of index word groups from a to-be-matched text, and generating one or more to-be-retrieved word groups; performing secret processing on the one or more to-be-retrieved word groups, and generating one or more ciphertext retrieval word groups; matching the one or more ciphertext retrieval word groups with ciphertext index word groups; when there is a ciphertext retrieval word group matching a ciphertext index word group, obtaining one or more keyword lengths; extracting continuous words satisfying the one or more keyword lengths from the to-be-matched text, and generating one or more to-be-matched word groups; performing secret processing on the one or more to-be-matched word groups, and generating one or more ciphertext matching word groups; and matching the one or more ciphertext matching word groups with ciphertext keywords, and obtaining a matching result. According to the method of the first aspect, the data processing amount occupied by the matching operation can be effectively reduced, and the execution efficiency of the matching operation is improved.
Owner:HUAWEI TECH CO LTD

Identifying root causes of test failures

A disclosed method defines root cause failure categories for a test case and associates each category with a corresponding test configuration property. Test configuration property information, indicative of the test configuration properties, are embedded in test case metadata. After generating test results, including test script messages, the test script messages assessed to identify the “N” most significant keywords in the message. The significance of a term in the test script message may be calculated based on an inverse document frequent parameter, independent of the term frequency within the document. Test result groups may then be determined by invoking a suitable clustering algorithm, e.g., a k-means clustering algorithm, to cluster the test script messages based on their corresponding keyword sets. Hypothesis test statistics may then be calculated for each test result group. A most probable root cause may then be identified for some or all of the test result groups.
Owner:DELL PROD LP

Machine learning techniques for context-based document classification

ActiveUS12675732B2Context basedData mining
Various embodiments of the present invention provide methods, apparatus, systems, computing devices, computing entities, and / or the like for performing context-based document classification prediction using a hierarchical attention-based keyword classifier machine learning framework. Certain embodiments of the present invention utilize systems, methods, and computer program products that perform context-based document classification prediction using at least one of techniques using contextual keyword classifications, techniques using attention-based keyword classifier machine learning framework, techniques using a greedy matching indicator, and / or the like.
Owner:UNITEDHEALTH GROUP INC

A multi-agent-based retrieval method, apparatus, device, and medium

The application relates to the technical field of artificial intelligence, in particular to a retrieval method and device based on multiple agents, equipment and a medium. Applied to a financial scene, the application realizes the intellectualization and dynamicization of the retrieval process through a multiple-agent cooperative working mechanism. The retrieval core elements are accurately extracted and a dynamic cognitive graph is updated by analyzing the agents, deep semantic mining and correlation expansion are carried out based on the cognitive seeds of semantic agents, the dependence on keywords in traditional retrieval is effectively broken, and the breadth and depth of semantic understanding are improved. The query statement is optimized by combining the extended semantics and the dynamic cognitive graph of the context agent, so that the query is more in line with the real intention and context of the user. The retrieval agent efficiently retrieves the candidate path in the updated dynamic cognitive graph, the recommendation agent filters the optimal result through cognitive consistency scoring, and the accuracy, relevance and logical coherence of the retrieval result are ensured, and the accuracy of the complex information retrieval task is significantly improved.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

A hybrid retrieval method and device, electronic equipment and storage medium

PendingCN122388194ASemantic searchEngineering
The application provides a hybrid retrieval method and device, electronic equipment and storage medium. By obtaining to-be-queried content, keyword retrieval and semantic retrieval are performed on the to-be-queried content based on a preset document database, a plurality of first candidate documents are obtained, a confidence weight is adjusted according to a preset weight adjustment strategy, a first confidence weight corresponding to keyword retrieval and a second confidence weight corresponding to semantic retrieval are obtained, a fusion score of each first candidate document is determined according to the first confidence weight and the second confidence weight, and a target document is determined from each first candidate document based on the fusion score. The advantages of keyword retrieval and semantic retrieval are combined, the accuracy and robustness of the retrieval result are improved, in addition, the confidence weights of keyword retrieval and semantic retrieval are adjusted, so that the retrieval strategy can adapt to the query requirements of the current scene, and the efficiency and quality of information retrieval are further improved.
Owner:CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD

A hallucination dataset generation method and device based on semantic confusion induction

PendingCN122154700ASemantic analysisInference methodsApplications of artificial intelligenceData set
The application provides a hallucination dataset generation method and device based on semantic confusion induction, which comprises the following steps: obtaining a plurality of hallucination induction instructions to be evaluated based on the keywords in the pre-constructed fact space conflict word set; calculating the semantic vulnerability index corresponding to each instruction based on semantic contradiction and context ambiguity, and performing instruction screening according to the semantic vulnerability index corresponding to each instruction to determine high-confidence hallucination induction instructions; inputting the high-confidence instructions into a large language model to make the model output text, and calculating the hallucination confidence level of the evaluation text by combining the hallucination probability index calculation method; storing the text with a hallucination confidence level exceeding a level threshold and the corresponding high-confidence hallucination induction instructions as a hallucination dataset. The application is a hallucination dataset generation method of systematic induction and quantitative detection, which is used to reveal the potential vulnerability of the model and support the safe application of artificial intelligence systems in complex contexts.
Owner:BEIJING UNIV OF POSTS & TELECOMM +1

A method and system for generating intellectual property legal facts

The application discloses a kind of generation method and system of intellectual property legal fact.The method includes: intellectual property evidence data acquisition step, according to the oral statement of voice recognition subsystem transcription, according to the written material scanned by text recognition subsystem, generate initial intellectual property evidence data set with time stamp and fixed key set;Intellectual property key word extraction step, based on the keyword list in the index library of multi-source legal corpus and synonym normalization rules, extract the candidate intellectual property key word containing subject, behavior, time, place, causality and right and obligation from the initial intellectual property evidence data set;Intellectual property legal fact generation step, based on timeline reconstruction and role consistency check, the candidate intellectual property key word is formed into intellectual property legal fact, while the corresponding relationship between the intellectual property legal fact and the initial intellectual property evidence data set is increased for retrieval and review.The method can improve the speed and accuracy of intellectual property evidence sorting, reduce irrelevant and ambiguous expression residues, and ensure that intellectual property legal fact and original evidence are one-to-one corresponding through traceable mapping, and enhance reviewability.
Owner:BEIJING FORESTRY UNIVERSITY

Information processing device

An information processing device includes an acquisition unit that acquires document data, a first extraction unit that extracts a first keyword from the document data in its entirety, a second extraction unit that extracts a second keyword from a text box included in the document data, a determining unit that finds a difference set between the first keyword and the second keyword, and determines a non-redundant keyword based on the difference set, and a classification unit that classifies the document data by assigning a classification tag to the document data. The classification unit excludes the classification tag related to the non-redundant keyword from candidates for the classification tag to be assigned to the document data, and then assigns the classification tag.
Owner:TOYOTA JIDOSHA KK

A training method of a source evaluation model, a source evaluation method, and related products

PendingCN122285891ARealize evaluationEvaluation resultMedicine
This application discloses a training method for a source evaluation model, a source evaluation method, and related products. Text sample sequences and frequency sample sequences are input into the evaluation model to be trained. The model encodes the text sample sequences and frequency sample sequences to obtain text sample vectors corresponding to the text sample sequences and frequency sample vectors corresponding to the frequency sample sequences. The evaluation model then performs prediction and evaluation processing on the text sample vectors and frequency sample vectors to obtain source evaluation prediction results. Based on the difference between the source evaluation result labels and the source evaluation prediction results, the parameters of the evaluation model are adjusted until the adjusted model meets the model training cutoff condition, and training ends to obtain the source evaluation model. Thus, this application can construct a source evaluation model based on the source sample name and the title, keywords, and publication frequency of the published sample text as features, thereby achieving the evaluation of the source.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

A scientific literature clustering method and system based on unsupervised keyword extraction

ActiveCN117453912BDegree of similarityHeadword
This invention relates to a scientific literature clustering method and system based on unsupervised keyword extraction. First, it effectively extracts keywords from scientific literature by comprehensively considering factors such as the occurrence of words in document abstracts and titles, the semantic similarity between words and the documents themselves, and the characteristics of domain keywords. Then, based on the characteristics of Chinese and English, this invention uses different embedding methods to cluster the extracted Chinese and English keywords, thereby achieving effective clustering of Chinese and English scientific literature. This invention considers the importance of words from multiple perspectives, comprehensively considering their occurrence in document abstracts and titles, and uses a method that automatically adjusts the preset keyword length based on domain characteristics to calculate keyword scores, thus incorporating more features of the words. This invention improves upon existing unsupervised keyword extraction algorithms.
Owner:SHANDONG UNIV

system

We provide the system. [Solution] A means for receiving and analyzing spatial design information input by the user, A means for generating a three-dimensional virtual environment design based on analyzed information, A means of visualizing the generated virtual environment design in a virtual domain and providing it to the user, A means of receiving feedback from users and using it to improve the accuracy of the generation method, A method for extracting important keywords and themes from spatial design information using natural language processing technology, A system that includes this.
Owner:SOFTBANK GROUP CORP

Document search method, device, storage medium, equipment and program product

The application discloses a document search method and device, a storage medium, equipment and a program product, and is applied to enterprise office, medical health record, online library and the like. The method comprises the following steps: in response to a search keyword input by a target object, at least one target text block is acquired, the content of the target text block comprises the search keyword, and the target object has a viewing permission of the target text block; a document name of at least one target document to which the at least one target text block belongs is displayed, and a context in which the search keyword is located in the at least one target text block is displayed; the target text block is stored in a resource library, the resource library comprises a plurality of first documents, each first document is composed of a plurality of first text blocks, each first text block corresponds to a record in the resource library, the record comprises content corresponding to the first text block and permission information, and the permission information is used for indicating an object who has a right to view the first text block. The application realizes the functions of improving information security and avoiding unauthorized access.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Complaint text classification method and device, computer device and storage medium

The application relates to a customer complaint text classification method and device, computer equipment and a storage medium. The method comprises the following steps: preprocessing text data to be processed to obtain a plurality of word segmentation results; performing vectorization processing on the plurality of word segmentation results to obtain first vector values of the word segmentation results; inputting the plurality of word segmentation results into a pre-trained topic analysis model to obtain second vector values of each customer complaint topic to which the text data to be processed belongs and third vector values of keywords of the text data to be processed; performing splicing processing on the first vector values, the second vector values and the third vector values to obtain splicing features, and taking the splicing features as features of the text data to be processed; and inputting the features of the text data to be processed into a topic classification model and a perception classification model of a pre-trained classification model respectively to obtain a customer complaint topic category and a customer complaint perception category to which the text data to be processed belongs. The method can deeply analyze specific customer complaint contents under a field category.
Owner:SHANGHAI PUDONG DEVELOPMENT BANK

A corpus processing method and system for large model search engines

PendingCN122332635ASingle sentenceSemantics
This application discloses a corpus processing method and system for a large-scale model search engine, relating to the field of data processing technology. The corpus processing method for a large-scale model search engine includes: obtaining a target word segmentation result based on a first sentence; obtaining the semantic contribution degree corresponding to each word element based on the target word segmentation result; filtering out word elements in the target word segmentation result whose semantic contribution degree is lower than a first preset value to obtain an optimized word segmentation result; obtaining the priority of the first sentence based on the optimized word segmentation result; and determining whether to input the first sentence as valid corpus into the large-scale model based on the priority. This application overcomes the limitations of traditional word segmentation and filtering, accurately identifying high-quality, low-frequency corpus with reliable sources and core semantics, preventing it from being overwhelmed by high-frequency, low-quality information, and simultaneously eliminating false, low-quality corpus based on keyword stuffing from the source, thereby improving the quality of the input corpus for the large-scale model.
Owner:ALIBABA TECHNOLOGY (GUANGZHOU) CO LTD