Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

78 results about "Trie" patented technology

In computer science, a trie, also called digital tree or prefix tree, is a kind of search tree—an ordered tree data structure used to store a dynamic set or associative array where the keys are usually strings. Unlike a binary search tree, no node in the tree stores the key associated with that node; instead, its position in the tree defines the key with which it is associated. All the descendants of a node have a common prefix of the string associated with that node, and the root is associated with the empty string. Keys tend to be associated with leaves, though some inner nodes may correspond to keys of interest. Hence, keys are not necessarily associated with every node. For the space-optimized presentation of prefix tree, see compact prefix tree.

Language model processing

Techniques for constraining a language model's generation / decoding process using a finite state machine are described. A finite state machine may represent information corresponding to APIs, arguments and argument values that are available / supported by the system. In some embodiments, a language model (LM) is constrained to generate tokens representing valid API calls based on the finite state machine. The system may enable generation of argument values from a defined set or free-form generation of argument values. The system may also enable unconstrained generation of a response by the LM. In some embodiments, a trie data structure is used to determine the possible next tokens that the LM can generate from.
Owner:AMAZON TECH INC

Large language model reasoning acceleration method and device based on two-stage speculative decoding and storage medium

The invention discloses a large language model reasoning acceleration method and device based on two-stage speculative decoding and a storage medium, and the method comprises the steps: constructing and initializing a Trie tree, and inserting a historical corpus, and phrase sequences in a document library or a code library into the Trie tree one by one; in the reasoning process, longest prefix matching is carried out based on a Trie tree, and a candidate draft sequence is generated by adopting branch backtracking and recursive search; performing confidence evaluation on the candidate draft sequence, calculating a joint confidence score of the sequence through probability multiplication and a Top-K screening mechanism, and judging whether the joint confidence score reaches a confidence threshold; if the accumulated confidence of the candidate sequence reaches a threshold value, skipping a small model generation stage, and directly entering large model verification; otherwise, entering a small model draft completion stage; and the final large model takes the replaced and updated draft sequence as final output. According to the method, adaptive acceleration of the decoding process can be realized, and the long text reasoning delay of the large language model is remarkably reduced while the generation quality is ensured.
Owner:ZHEJIANG UNIV

Generating model output using a knowledge graph

Techniques for constraining the results of a generative language model to valid information using knowledge-grounded documentation. A generative language model may generate invalid results, including compound entities and incorrect entity relations. The techniques include, for a given user inquiry, determining a set of documented information, from a particular knowledge base, that corresponds to the user inquiry. The techniques further include determining a subgraph from a knowledge graph representing the knowledge base, as well as determining a trie data structure representation of the set of documented information. The user inquiry and subgraph are provided as input to a trained generative language model for generating a response to the user inquiry. The techniques include using the trie data structure to validate that the generated response corresponds to real information from the set of documented information.
Owner:AMAZON TECH INC

Encoding and decoding TRIE data structures for enhanced generation of customer journey analytics

A method for efficiently encoding a trie data structure for transmission according to an embodiment includes receiving an application programming interface (API) request pertaining to the trie data structure that is indicative of flows of customer interactions with automated agents of a contact center, obtaining the trie data structure in which each of multiple nodes has an associated prefix key that defines a path from a root to the corresponding node, and encoding the nodes in a transmission format having a dictionary data structure. The nodes in the transmission format do not have the prefix key that defines the path from the root to the corresponding node. The method also includes transmitting a response to the API request based on the trie data structure encoded in the transmission format.
Owner:GENESYS CLOUD SERVICES INC

File anti-desensitization self-learning recognition system and method based on information entropy

The invention discloses a file anti-desensitization self-learning recognition system and method based on information entropy, belongs to the technical field of intersection of natural language processing and content security recognition, and is applied to document screening and risk recognition in a multi-task scene. The implementation method comprises the following steps of: 1, performing character recognition and noise reduction processing on an original file to form a data set; 2, training labeled sample data through small samples, respectively adopting probability distribution of a data sliding window and information entropy to carry out maximum and minimum normalization screening, and further utilizing a fitted linear regression model to form an anti-desensitization word list; 3, screening the anti-desensitization degrees of the sentence segments of the data set by adopting a dictionary tree Trie structure to form an anti-desensitization sentence segment table; 4, marking the chapter-level anti-desensitization degree data set text fragments by using the large model; 5, generating an anti-desensitization report according to the anti-desensitization word and the anti-desensitization degree of the marked anti-desensitization file; compared with the prior art, the anti-desensitization file screening method and device have the advantage that the anti-desensitization file screening accuracy is improved.
Owner:BEIJING INST OF TECH

Multi-modal fusion short message compliance and security dual-auditing method and multi-modal fusion short message compliance and security dual-auditing system

The invention discloses a multi-modal fusion short message compliance and security dual auditing method and system, and the method comprises the steps: encrypting account information through employing an irreversible algorithm based on SHA256 and a random salt value, and decomposing the content of a short message into a text stream, a link stream and a symbol stream through employing a regular expression; traversing each character in the short message content by adopting a prefix mode based on a Trie tree in combination with an AC automaton algorithm to detect sensitive words; performing symbol semantic classification mapping and analysis on the special symbol feature data extracted from the symbol stream; carrying out sending behavior analysis on the text feature data extracted from the text stream; performing special detection on link feature data extracted from the link stream; based on the factors corresponding to the sensitive words, the semantics, the behaviors and the links and the weights of the factors, a multi-modal feature fusion decision risk assessment algorithm is adopted to output a risk score value and a risk decision rule of the risk score value so as to execute short message interception operation or short message release operation. According to the invention, full-dimension perception and dynamic defense of risks can be realized.
Owner:JIANGXI TIANLI TECH INC

Combined word matching detection method and system, electronic device and storage medium

The present application relates to a combined word matching detection method and system and a medium. The method comprises: acquiring a preset first combined word, splitting the first combined word to obtain a plurality of first word elements, and, according to the first word elements and the first combined word, constructing a first relationship mapping table; according to the first relationship mapping table, generating a trie; acquiring a text to be subjected to detection, performing, according to the trie, word element matching on said text to obtain a matching record list, and performing aggregation processing on the matching record list to obtain a combined word index dictionary; determining a word element minimum interval of the combined word index dictionary, and, according to the word element minimum interval, calculating a matching score of each combined word in the combined word index dictionary, so as to obtain a matching score list; and, according to the text length and the combined word length, performing normalization processing on the matching score list, so as to obtain a combined word matching result of said text.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

IP address matching method and device

An IP address matching method and device, the method being applied to a routing forwarding device, the method comprising: constructing a multi-bit Trie data structure in the routing forwarding device, the odd layer nodes of the multi-bit Trie being odd nodes, and the even layer nodes of the multi-bit Trie being even nodes; the index information of the Chinese child nodes, the index information of the child nodes and the next-hop information are stored through the odd nodes, and the next-hop information is stored through the even nodes; according to the bit sequence of the target IP address, layer-by-layer searching is started from a root odd node, and whether Chinese daughter nodes exist or not is judged based on a Chinese daughter bitmap of the odd node; skipping the even node and directly skipping to the next odd node under the condition that the Chinese child node exists; and under the condition that the Chinese child nodes do not exist, checking the child bitmaps of the odd nodes to position the even nodes, and obtaining the longest matched next hop information in the even nodes. According to the method, the access and calculation overhead of the child nodes can be saved, and the equivalent search step length is increased.
Owner:COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI

Method and system for identifying finance and tax question and answer sensitive information based on large model

The invention discloses a large model-based finance and tax question and answer sensitive information identification method and system, and the method comprises the steps: obtaining finance and tax question data, and carrying out the processing of the obtained finance and tax question data; key prohibited word detection is carried out on the processed finance and tax question data through an established Chinese pinyin sensitive word Trie tree; when it is judged that no key forbidden word exists in the finance and taxation question data, detecting violation sensitive words through violation semantics in a trained finance and taxation large model; and when no violation sensitive word is detected in the output of the large finance and taxation model, outputting the finance and taxation question data to a normal question and answer system. According to the method, finance and taxation violation oriented questions and answers are screened through the finance and taxation large model obtained through training, the Chinese pinyin sensitive word Trie tree is constructed, the detection effect is improved by training the semantic comprehension ability of the large model, and sensitive text recognition is achieved.
Owner:AISINO CORPORATION

Efficient processing of TRIE data structures to support customer journey visualizations

A method for providing efficient trie data structure processing according to an embodiment includes splitting, based on organization identifiers and sequence identifiers, a data frame indicative of a set of events associated with customer interactions with automated agents of a contact center to produce a set of multiple partitions, and producing a set of multiple trie data structures, including generating a trie data structure for each partition. Each trie data structure represents aggregate counts of a corresponding subset of event sequences associated with a corresponding organization. The method also includes merging multiple trie data structures of the set of trie data structures to produce a combined organization trie data structure and storing the combined organization trie data structure to enable generation of a visualization of the combined organization trie data structure.
Owner:GENESYS CLOUD SERVICES INC

Knowledge enhancement-based non-auxiliary model speculation reasoning method

The invention discloses an auxiliary-model-free speculation reasoning method based on knowledge enhancement. The method comprises the steps that firstly, a lightweight external knowledge base is constructed based on a segmented caching strategy; secondly, constructing a multi-branch prefix tree Trie supporting dynamic path prediction and context alignment based on an input token sequence by adopting a step-by-step insertion and structural node multiplexing mechanism; generating a tree candidate token path based on the Trie tree, generating a library candidate token path based on an external knowledge base, combining and de-duplicating to generate a mixed candidate token path, constructing a multi-branch tree sharing a prefix, performing parallel reasoning on all nodes in the multi-branch tree through a large model, verifying and updating the Trie tree, and circularly executing the operation to complete speculative reasoning. According to the method, the reasoning performance in a scene without an auxiliary model is comprehensively improved, and a more efficient and more stable generation capability is provided for large model reasoning in tasks in the fields of medical treatment, law and the like.
Owner:HANGZHOU DIANZI UNIV

MPT-based block chain database system, data access method, terminal and medium

The invention relates to the field of data access, and particularly provides an MPT-based block chain database system, a data access method, a terminal and a medium, the system comprises a native MPT storage engine used for realizing an MPT node structure in a disk and a memory, and the node structure comprises branch nodes, extension nodes and leaf nodes; the asynchronous I / O module is used for realizing non-blocking disk operation based on Linux iouring; the file system bypass module is used for supporting direct access to the block device; the versioning concurrency control module is used for realizing lock-free reading through an immutable Trie structure and atomic pointer updating; and the dynamic compression module is used for dynamically adjusting a historical data retention strategy according to the disk space utilization rate. The MPT structure storage engine is realized, basic data support is provided for a high-performance block chain system, serialization overhead is eliminated, and storage efficiency is improved.
Owner:INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD

Method, computer readable medium and computer system for implementing a trie data structure with a sub-trie tree data structure

This application discloses techniques related to tree data structures capable of storing information indicative of database key codes. A computer system can operate a database. The computer system can store a multi-level tree data structure capable of being used to perform key code lookups for the database. In various cases, the multi-level tree data structure can be stored in system memory as a plurality of sub-tree data structures, each sub-tree data structure comprising a set of linked nodes. A given one of the plurality of sub-tree data structures can be stored in system memory as a respective contiguous block of information. The computer system can access the respective contiguous block of a first particular sub-tree data structure encompassing a particular range of levels in the multi-level tree data structure. The access can be performed without accessing one or more other sub-tree data structures encompassing one or more levels within the particular range of levels.
Owner:SALESFORCE INC

Efficient Use of TRIE Data Structure in a Database

The present invention provides a time-saving method for performing queries in a database or information retrieval system, the method comprising performing operations such as intersection, union, difference, and exclusive OR on two or more key sets stored in the database or information retrieval system. In a novel execution model, all data sources and operations are tries, and the input tries for higher-order set operations can be the output of lower-order set operations that are evaluated on demand. Two or more input tries are combined according to the corresponding set operations to obtain the key sets associated with the nodes of the corresponding result tries. The physical algebra of the bitmap-based trie implementation directly corresponds to the logical algebra of the set operations and allows for an efficient implementation by means of bitwise Boolean operations.
Owner:CENSHARE GMBH

Look ahead strategy for trie-based beam search in generative retrieval

Systems and methods are provided for generating a keyword sequence from an input query. A first text sequence corresponding to an input query may be received and encoded into a source sequence representation using an encoder of a machine learning model. A keyword sentence may then be generated from the source sequence representation using a decoder of the machine learning model. The decoder may generate a modified generation score for a plurality of prediction tokens, wherein the modified generation score is based on the respective prediction token generation score and a maximum generation score for a suffix of each prediction token. The decoder may then select the prediction token of the plurality of prediction tokens based on the modified generation score, and add the selected prediction token to the previously decoded partial hypothesis provided by the decoder.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

An efficient spatial keyword query method based on association rule mining

The application relates to an efficient spatial keyword query method based on association rule mining, and belongs to the field of spatial keyword query. The application comprises a data preprocessing stage, an association rule mining stage, an index construction stage and a spatial keyword query stage. The association rule mining stage selects corresponding frequent item sets based on a depth control materialization strategy. The index construction stage constructs a quadtree index for the spatial part of a data set, and creates a trie and an inverted list for the text part of the data set. The query stage adopts a coarse-grained spatial query, and combines the inverted list and the corresponding materialized inverted list of the frequent item sets to perform retrieval. The application combines association rule mining, materialized inverted lists and spatial keyword query, and can greatly improve the query efficiency of spatial keywords.
Owner:YUNNAN NORMAL UNIV

A method and system for detecting prefix hijacking for uncertain information sources

The present invention provides a method and system for detecting prefix hijacking for uncertain sources. The method comprises: obtaining data from the routing information base, resource public key infrastructure, and Internet routing registry in the Border Gateway Protocol; constructing prefix cloud droplets from the time, space, and data source dimensions based on cloud model theory, and deriving deterministic data and uncertain data based on the discreteness of the cloud model; performing spatiotemporal stability analysis on the deterministic data, dynamically scoring each prefix cloud droplet, and establishing a dynamic deterministic information base; for the uncertain data, performing online route detection using a custom dictionary tree; utilizing prefix coverage and matching rules in combination with a dual verification mechanism to determine whether the route is legitimate, thereby completing anomaly detection and online updating. The present invention proposes a prefix hijacking detection method based on cloud model theory and a Trie tree to overcome the characteristics of the Border Gateway Protocol, such as dynamic changes, complex and large-scale networks, and uncertain data.
Owner:NANJING UNIV OF POSTS & TELECOMM

Elastic service optimization method for block chain data storage

The invention belongs to the technical field of block chains, and particularly relates to an elastic service optimization method for block chain data storage, which effectively reduces the storage cost of a block chain by optimizing the storage mode of account book data and state data and designing a network maintenance scheme adapted to the dynamic nature of a permission removal network, and improves the efficiency of the block chain. And the data availability and the system stability are improved. According to the method, storage optimization is carried out on account book data by adopting batch coding and a height-based coding method; a dual Trie state management system is introduced into state data, and technologies of state expiration, mining, creation and the like are designed to realize efficient management; meanwhile, dynamic changes of the nodes are coped with through a group upgrading and degrading mechanism and a block updating method.
Owner:SHANDONG UNIV

A construction method of a superset index structure combining TRIE and LOUDS

The present invention relates to a method for constructing a superset index structure combining TRIE and LOUDS, belonging to the technical field of set and string processing. The present invention includes a data preprocessing stage, an index structure construction stage, and a superset query stage. In the data preprocessing stage, the sets and elements in the original set dataset are mapped and sorted. In the index structure construction stage, a hybrid index structure with TRIE at the upper part and LOUDS at the lower part is constructed. In the superset query stage, a query is given, and all sets that are subsets of the given query are retrieved on the constructed hybrid index structure. The present invention can make full use of the high query efficiency of TRIE and the high space compressibility of LOUDS, enabling the frequently accessed upper part to have a fast query speed and the less frequently accessed lower part to have high compression performance.
Owner:YUNNAN NORMAL UNIV

A text key phrase extraction method, storage medium and device integrating inter-sentence correlation relationship

The application belongs to the field of natural language processing, and particularly relates to a text key phrase extraction method, a storage medium and a device that integrate inter-sentence correlation, comprising: extracting a nominal phrase in a part-of-speech combination mode; filtering the nominal phrase by combining two Trie trees to obtain a candidate phrase set; calculating a global semantic similarity score of the candidate phrase; clustering each sentence to obtain a sentence cluster containing different semantic information; and sorting the candidate phrases in the sentence cluster according to the global semantic similarity score of the candidate phrase to obtain a key phrase set; the application integrates the inter-sentence relationship in the unsupervised key phrase extraction model based on embedding, improves the accuracy of the model, and constructs different Trie trees to calculate mutual information and left and right information entropy as a candidate phrase extraction method in the key phrase extraction model, thereby reducing the probability of incomplete semantic information appearing in the extracted candidate phrase set.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Inference methods for word or wordpiece tokenization

Systems and methods for performing inference for word or wordpiece tokenization are disclosed using a left-to-right longest-match-first greedy process. In some examples, the vocabulary may be organized into a trie structure in which each node includes a precomputed token or token_ID and a fail link, so that the tokenizer can parse the trie in a single pass to generate a list of only those tokens or token_IDs that correspond to the longest matching vocabulary entries in the sample string, without the need for backtracking. In some examples, the vocabulary may be organized into a trie in which each node has a fail link, and any node that would share token(s) or token_ID(s) of a preceding node is instead given a prev_match link that points back to a chain of nodes with those token(s) or token_ID(s).
Owner:GOOGLE LLC

High-load micro-service starting method and system, electronic equipment and storage medium

The invention relates to the technical field of computers, and discloses a high-load micro-service starting method and system, electronic equipment and a storage medium, and the method comprises the steps: executing a hierarchical preloading mechanism, initializing a connection pool, and building a preset minimum idle connection number; checking algorithm optimization is executed, a client feature Trie tree is constructed, root nodes are initialized, child nodes are inserted in sequence according to character string characters, new nodes are created when the child nodes are missing, and the nodes are marked as word ending when character strings are ended; executing request verification mechanism optimization, and returning an error response or an unauthorized response when verification fails; probe parameter dynamic adjustment is executed, and the survival probe detection interval is dynamically adjusted according to the operation indexes; according to the method and the device, the problems of resource waste, low processing efficiency and insufficient stability in a traditional scheme are solved, and reliable technical guarantee is provided for a high-load micro-service scene.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Encoding and decoding method and system for route origin authorization (ROA)

Described are an encoding method and system for ROAs. The encoding method includes the following steps: given a set of authorized IP prefixes an AS which are maintained with an IP address trie. By specifying a sequence of hanging levels on the IP address trie, it is divided into a set of non-overlapping sub-trees, each rooted at a hanging level. A node on a hanging level uniquely defines a sub-tree rooted at it, whose prefix can be encoded as the identifier of this sub-tree. All authorized prefixes covered by a sub-tree can be encoded into a bitmap of 2h bits, where h is the height of this sub-tree. Thus, the set of authorized IP prefixes of an AS is encoded into several tuples (identifier, bitmap).
Owner:COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI

Word completion method and apparatus

Embodiments of the present application disclose a word completion method applied to a search scenario to complete an incomplete word input by a user. The method of the embodiments of the present application is based on an improved dictionary tree, hot words are stored in part nodes of the dictionary tree, in the word completion method, a target node matching the string is searched in the dictionary tree Trie, and at least one completed word is output to the user based on the hot words stored in the target node. The word completion efficiency can be improved, and the user is prevented from being recommended words when inputting a too-short string.
Owner:HUAWEI TECH CO LTD

A symbol fast retrieval method for super-large multi-station PLC project

PendingCN122332617ADifferential codingTrie
This application relates to the field of PLC programming technology, and in particular to a fast symbol retrieval method for ultra-large multi-station PLC projects. The method involves data parsing, symbol metadata standardization, and fragment preprocessing of the PLC project file to obtain standardized SymbolMeta fragment data. Based on the standardized SymbolMeta fragment data, a three-level hierarchical index architecture is constructed, consisting of station-level hash routing, task-level compressed prefix Trie trees, and symbol-level differential encoded inverted lists. According to user query conditions, station-level fast pruning, task-level Trie tree matching, and symbol-level differential decoding aggregation are performed sequentially within the three-level hierarchical index architecture to generate structured retrieval results. Incremental updates are performed on the three-level hierarchical index architecture based on the PLC project's modification instructions. This solves the problems of high latency, weak fuzzy matching capability, high resource consumption, and delayed updates in retrieval of symbols exceeding 100,000, achieving millisecond-level accurate positioning.
Owner:CLP INTELLIGENT TECH CO LTD

Semantic identification blockchain retrieval method and system for large-scale distributed data

This invention belongs to the field of blockchain technology, specifically relating to a semantic identifier blockchain retrieval method and system for large-scale distributed data. This invention designs a unified semantic identifier model, mapping data records from different sources and in different formats to multi-level scalable semantic paths. It constructs a MIR tree forest, integrating the organization of Trie and the verifiability of Merkle, with each scenario as the root. Only the root hash of each scenario or its aggregate root is maintained in the block header, thereby achieving overall constraint on the state of large-scale off-chain semantic indexes while maintaining lightweight on-chain storage overhead. Simultaneously, this invention designs corresponding index traversal and verification object construction mechanisms around various semantic retrieval modes, enabling clients to verify the correctness of target records with only limited path and node hash information. This provides a new technical path for unified semantic management and efficient verifiable retrieval of multi-source heterogeneous data in the HSB environment.
Owner:ZHEJIANG SCI-TECH UNIV +1

Method for quickly loading ten-million-level domain name rules based on distributed compiling

The invention discloses a ten-million-level domain name rule quick loading method based on distributed compiling, which comprises the following steps of: reading and analyzing a rule file containing an accurate domain name and an extensive domain name, and distributing rule analysis result data into a hash table according to a grouping number to obtain a plurality of rule groups; according to the grouping number, compiling threads with the same number are created to form a compiling thread pool; traversing all the rule groups, allocating a compiling thread for each rule group to execute a compiling task, and performing distributed parallel compiling by using a Hyperscan state machine and a Trie tree; receiving an input domain name character string and determining a rule group to which the input domain name character string belongs; if yes, the Hyperscan state machine of the rule group is used for preliminary matching, and if yes, the Trie tree of the rule group is used for accurate matching. According to the method, the problem that a Hyperscan lightweight interface cannot precisely limit a domain name is solved, and rapid loading of a ten-million-level domain name rule is realized.
Owner:北京九栖科技有限责任公司

IP address searching method and device

An IP address searching method is applied to IPv6 and comprises the steps that a searching data structure tree is built based on a multi-bit Trie tree, and each Trie node internally provided with a real node corresponds to one or more IP prefixes in an IP forwarding table; determining a mappable pipeline level range of each node according to an inverse distance of each node in the data structure tree, a pipeline level where a child node is located, a path length from the node to a root node and a position of a pipeline level where a father node is located; wherein the inverse distance is defined as the maximum distance between the node and all subsequent leaf nodes; dynamically mapping each node into an independent storage resource of the corresponding pipeline level according to the mappable pipeline level range and the storage space condition of each pipeline level; and in each pipeline level, parallelly executing IP address searching according to a node mapping result so as to realize pipeline parallel searching. The method can improve the searching efficiency.
Owner:COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI

Prefix hijacking detection method and system for uncertain information source

The invention provides an uncertain information source-oriented prefix hijacking detection method and system. The method comprises the following steps of: acquiring a routing information base, a resource public key infrastructure and Internet routing registration center data in a border gateway protocol; prefix cloud droplets are constructed from time, space and data source dimensions based on a cloud model theory, and deterministic data and uncertain data are obtained according to the dispersion degree of a cloud model; performing space-time stability analysis on the deterministic data, dynamically scoring each prefix cloud droplet, and establishing a dynamic deterministic information base; for uncertain data, a self-defined dictionary tree is used for carrying out online detection on a route; and judging whether the routing is legal or not by utilizing prefix coverage and matching rules and combining a dual verification mechanism, so as to complete anomaly detection and online updating. The invention provides a prefix hijacking detection method based on a cloud model theory and a Trie tree so as to overcome the characteristics of dynamic change of a border gateway protocol, complex network, large scale, data uncertainty and the like.
Owner:NANJING UNIV OF POSTS & TELECOMM

A sensitive word detection method and device, computer equipment and a storage medium

The embodiment of the application belongs to the technical field of natural language processing, and relates to a sensitive word detection method and device, computer equipment and a storage medium. The method comprises the following steps: receiving a sensitive word detection request sent by a user terminal; calling a Chinese character Unicode encoding white list, and performing an interference character clearing operation on input text data according to the Chinese character Unicode encoding white list to obtain pure text data; performing a character merging operation on the pure text data to obtain to-be-matched text data; calling a sensitive word Trie tree, and performing a sensitive word similarity matching operation in a similarity matching chain between the characters of the to-be-matched text data and the root node character of the sensitive word Trie tree as a starting point to obtain matching result data; performing a comprehensive scoring operation on the matching result data to obtain a comprehensive scoring result; and outputting the comprehensive scoring result to the user terminal. The sensitive words in the user input text are automatically and efficiently identified and filtered, so that the safety and compliance of the system are protected.
Owner:ASPIRE INFORMATION TECH BEIJING